Skip to content

AI strategy & innovation

AI strategy that ends in a shipped use case, not a deck

We build the three things that decide whether AI spend returns anything: a measured baseline, a use-case portfolio with a written kill criterion per case, and an EU AI Act position that matches what you actually run. We will not open with an enterprise-wide data maturity programme — readiness is assessed only for the critical data elements your first three use cases consume, because the enterprise version takes a year and produces a radar chart.

30 minutes with the engineer who would do the work — not a salesperson. No obligation, and you keep whatever we work out on the call.

A working session scoring candidate AI use cases against a measured process baseline

What we design against

95%
of GenAI pilots return nothing measurable to the P&LMIT Project NANDA, "The GenAI Divide: State of AI in Business 2025" — 52 structured interviews, 153 survey leaders, 300+ public deployments. One non-peer-reviewed report, and four figures on this page come from it, which is why we name its weaknesses where we use it. Not our number; the reason our engagements are gated.
5%
of evaluated enterprise GenAI tools reach productionNANDA funnel: ~60% investigated, ~20% piloted, ~5% implemented. Generic chatbots convert at ~83%, so the failure sits in custom and vendor enterprise builds, not in the models.
8.55%
of Bulgarian enterprises use AI, against a 19.95% EU averageEurostat, annual EU survey on ICT usage and e-commerce in enterprises, 2025 reference year. By size across the EU: large 55.03%, medium 30.36%, small 17%. Bulgarian mid-market adoption is roughly half the EU average.
≤ 12 months
modelled payback we design first-wave use cases againstOur design target, not a measured client result. Anything projecting past 18 months in wave one is usually a strategic bet mislabelled as an efficiency case, and we score it as one.

What we build

Consultants and a client reviewing a printed data inventory

Readiness scoped to the data your use cases actually consume

Enterprise-wide data maturity assessments take a year and end in a radar chart. We assess only the critical data elements your candidate use cases read — usually 5 to 15 fields — profiling the real tables with dbt tests, Great Expectations or Soda rather than accepting a data dictionary. Entity resolution on customer, supplier and product master data is scored separately, because that is where back-office cases actually break. We have run that problem at scale: the European car marketplace we built carries 300,000+ live listings ingested from dealer feeds with inconsistent schemas, currencies and vehicle taxonomies, which is the point at which deduplication and entity resolution stop being theory.

  • A go / go-with-remediation / blocked verdict per use case, with the remediation priced in days rather than described as a programme
  • A critical-data-element register: named owner, system of record, lineage to the consuming process, measured completeness and duplication rate, freshness against the decision cadence
  • An integration surface note per system — which APIs exist, what permission model governs them, whether a non-production environment exists at all
Hand placing a note on a prioritisation grid of sticky notes

A portfolio with published weights and a kill criterion per case

Forty workshop ideas ranked by enthusiasm is not a portfolio. We score with published weights — value 35–45%, feasibility 25–35%, time-to-value 10–15%, reuse 10–15%, strategic fit 5–10% — with risk and adoption applied as penalties. We fix the exact weights with you in writing before any case is scored, they sum to 100, and we do not change them after seeing the scores; if a sponsor wants a re-weighting it happens in the open and the whole portfolio is re-ranked, including the cases it demotes. Value is never a 1–5 opinion: annual volume × measured handling minutes × fully loaded hourly cost × an automation rate bounded by your observed exception rate. Baselines come from Celonis, SAP Signavio or a two-week time study.

  • A ranked wave one of three to six cases with modelled payback under twelve months, each with a named accountable owner
  • The rejected list with the reason recorded — a healthy first gate retires 30–50% of candidates, and a 0% kill rate means the gate is decorative
  • A written kill criterion and a decommission condition attached to every case before any build starts
Compliance officer marking a passage in a printed regulation

EU AI Act classification, inventory and registration posture

The Digital Omnibus on AI, which amends Regulation (EU) 2024/1689, has applied since 27 July 2026 and changed 42 articles and 3 annexes: standalone Annex III duties now start 2 December 2027, Annex I embedded systems 2 August 2028, Article 50 transparency marking 2 December 2026. What did not move — Article 5 prohibitions, Article 4 AI literacy, GPAI obligations. We classify what you actually run, including AI switched on by default inside Dynamics 365, Salesforce, ServiceNow and the HR stack. Where a governance platform is on the table — Credo AI, Holistic AI, watsonx.governance, OneTrust — we test it against one question: does it replace a spreadsheet, or add a second one? Most mid-market clients need the register and the intake gate, not the platform. Regulatory position last checked August 2026; we re-check it quarterly, and every date we hand you carries its Official Journal citation in the deliverable rather than on a slide.

  • A classified AI register with the provider/deployer determination recorded and one named accountable owner per system, not a committee
  • Article 6(3) derogations written against all four conditions — and a system self-assessed as not high-risk must still be registered by its provider under Article 49(2). If you substantially modify a vendor system or run it outside its declared intended purpose, that provider is you. The filing is public and searchable, except law-enforcement, migration, asylum and border-control entries, which sit in the non-public section of the database under Article 49(4).
  • Deployer duties made concrete: Article 26(6) six-month log retention configured in the actual platform, Article 26(7) worker and works-council notification, Article 27 FRIA where it triggers, Article 14 human oversight assigned to a competent named person rather than a role, Article 13 instructions for use treated as the document your own compliance inherits from, and the Article 73 serious-incident trigger, escalation path and clock written before go-live rather than after the first one
  • If you already hold ISO 27001, ISO/IEC 42001 Annex A mapped onto the controls that certification already satisfies, with only the genuine gaps filled — not a parallel governance structure standing beside the one you audit every year
Colleagues weighing two options drawn in columns on a whiteboard

Build, buy or partner, with human review priced as a first-class line

A business case with one line for "AI costs" is not a business case. We model three options over 36 months with identical line items: token spend at measured volumes with model routing and semantic caching modelled explicitly, self-hosted GPU economics at real utilisation, evaluation harness build and upkeep, integration and permissions work, model migration effort, and human review — routinely 5 to 50 times inference cost and routinely omitted.

  • A 36-month comparison sensitivity-tested against a 10× fall in API pricing, not 50%. The Stanford AI Index tracked the cheapest model reaching GPT-3.5-level MMLU falling from roughly $20 to $0.07 per million tokens between late 2022 and late 2024 — that is not a price cut on a fixed model, it is capability getting cheaper to buy, and it is the assumption that most often invalidates a self-hosting case in month nine. We carry the series to the current quarter and stamp the model with its as-of date.
  • The self-hosting threshold stated as a volume rather than a vibe: on our reference build — 4×H100 SXM5 serving a 70B at FP8 via vLLM, priced at August 2026 EU rates — self-hosting starts to win somewhere above 20 to 50 million output tokens a day at sustained 70%+ utilisation. We compute your crossover on your measured traffic mix, and we model input and output tokens separately, because API pricing charges them at 3 to 5 times different rates and blending them is how self-hosting business cases get talked into existence.
  • A priced exit: accumulated prompts, fine-tunes, evaluation sets and integrations quantified as switching cost before you sign, not after
Engineer reviewing a table of test results on a wide monitor

An evaluation harness that makes the go/no-go falsifiable

A demo is not a measurement. We build a golden dataset from your own historical cases, deliberately over-sampling the hard tail — the exceptions, the ambiguous items, the Bulgarian-language documents and Cyrillic scans — labelled by your subject-matter experts rather than by us. Bulgarian and Cyrillic handling is not a claim we ask you to take on trust: the digital dictionary we built for the Ministry of Education and Science runs at beron.mon.bg and you can open it now. Retrieval recall@k is measured separately from generation, because most reported hallucination is a chunking or retrieval failure and is fixable once it is visible.

  • Per-field precision and recall for extraction, groundedness and citation accuracy for RAG. Working gates, tuned to your risk tier: recall@5 above 80% before any generation work is worth doing, groundedness 90–97%, per-field recall set field by field rather than blended into one number. Expect the first honest measurement to land 15–30 points below the demo impression — that gap is the point of measuring.
  • A regression suite that runs on every prompt, model, chunking or index change: golden set and traces in Langfuse or Arize Phoenix, scored with Ragas or DeepEval from pytest in CI, with Evidently watching input drift in production. Three layers, not one tool.
  • Where scoring is scaled past the golden set with an LLM judge, the judge is calibrated against human labels first and the inter-annotator agreement is stated — an uncalibrated judge is a confident number with nothing underneath it
  • Acceptance thresholds published and signed by the process owner before the pilot starts, with production traces instrumented to the OpenTelemetry GenAI semantic conventions
Developer studying a system diagram beside a hand-drawn sketch

Agentic scope, unit economics and a kill switch

Much of what is sold as agentic is RPA with a language model in front, and Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027 on escalating cost, unclear business value and inadequate risk controls. The version that survives is narrow: one workflow, a bounded tool set, the agent proposing while a workflow engine such as Temporal executes the side effects with retries, idempotency and an audit trail. Nothing irreversible happens without approval, and the fallback — what the process does when the agent is off, and how fast you revert — is written before the first tool call.

  • A per-task spend cap rather than a per-token budget, because agent loops fail expensively and silently — cost per completed task is the number we report
  • Failure modes tested deliberately instead of discovered in production: tool-call loops, stale context, cascading errors across steps, and prompt injection through retrieved content, which is the one your network controls never see
  • An AI Act position taken early — an agent acting in employment, credit or essential-services contexts sits in Annex III regardless of how the vendor brands it — and a decommission criterion written at the start rather than negotiated at month nine

Systems we work with

We integrate with what you already run. If a platform below is missing, tell us — the pattern usually transfers.

Regulation and standards we work against

  • Regulation (EU) 2024/1689 (EU AI Act) and the Digital Omnibus on AI
  • GDPR and EDPB Opinion 28/2024 on AI models and personal data
  • NIS2, DORA, the Data Act and the Data Governance Act
  • ISO/IEC 42001 Annex A (A.2–A.10)
  • ISO/IEC 23894
  • ISO/IEC 42005
  • ISO/IEC 5259
  • ISO/IEC 27001 as the ISMS we extend rather than duplicate
  • NIST AI RMF 1.0 and the Generative AI Profile (NIST AI 600-1)
  • OWASP Top 10 for LLM Applications
  • CEN-CENELEC JTC 21 harmonised standards work

Governance, evaluation and observability

  • Credo AI
  • Holistic AI
  • IBM watsonx.governance
  • OneTrust AI Governance
  • ServiceNow AI Control Tower
  • Collibra AI Governance
  • Langfuse
  • LangSmith
  • Arize Phoenix
  • Ragas
  • DeepEval
  • Evidently
  • OpenTelemetry GenAI semantic conventions

Data, platform and process evidence

  • Databricks Unity Catalog
  • Snowflake Horizon
  • Microsoft Purview
  • Apache Iceberg
  • dbt
  • Great Expectations
  • Soda
  • Airflow
  • Dagster
  • Celonis
  • SAP Signavio
  • UiPath Process Mining
  • Power Automate Process Mining
  • pgvector on PostgreSQL
  • Elasticsearch
  • Azure AI Search
  • Qdrant

Systems of record, orchestration and hosting we scope against

  • SAP S/4HANA
  • Microsoft Dynamics 365
  • Oracle Fusion
  • ServiceNow
  • Salesforce
  • Zendesk
  • Azure AI Foundry
  • AWS Bedrock and SageMaker
  • Google Vertex AI
  • Snowflake Cortex
  • Temporal, LangGraph and the Model Context Protocol for scoped agent work
  • vLLM, NVIDIA Triton and TensorRT-LLM for self-hosted serving
  • OVHcloud, Scaleway, Hetzner and StackIT for EU residency; INSAIT BgGPT for Bulgarian-language work

How the work runs

  1. 01

    Baseline and inventory

    2 weeks

    What you keep

    A Baseline Ledger — measured cost per transaction, cycle time, touchless rate and exception rate for the candidate processes — plus a classified AI register covering vendor features switched on by default and disclosed personal-account use.

    The decision it forces

    Which processes have a baseline solid enough to be scored, and which systems already carry obligations under Article 5, Annex III or Annex I.

    When we stop

    If no process owner will release transaction data or sign off a time study, we stop and hand you the ledger and the register. A portfolio scored on managers’ estimates is a deck with arithmetic on it.

  2. 02

    Portfolio scoring and value case

    2–3 weeks

    What you keep

    A scored portfolio with the weights fixed in writing before scoring, a per-case value model built from volume, measured handling minutes, loaded cost and an exception-bounded automation rate, and a written kill criterion per case.

    The decision it forces

    The three to six wave-one cases with modelled payback under twelve months, and the recorded reason every rejected case was rejected.

    When we stop

    If nothing clears a twelve-month modelled payback once human review cost is included, we say so and recommend upstream data or workflow fixes instead of a build.

  3. 03

    Architecture, TCO and AI Act position

    3–4 weeks

    What you keep

    A 36-month build/buy/partner model with identical line items across all three options, a routed architecture spec with the data-residency boundary drawn, and a remediation plan dated against the post-Omnibus deadlines.

    The decision it forces

    Where each workload runs, what it costs per successful task, and which obligations attach before go-live rather than after.

    When we stop

    If the routed design cannot keep regulated personal data inside the boundary the CPDP — the Bulgarian data protection authority — would accept, we hand over the architecture model and stop. We do not bill a pilot phase we cannot responsibly recommend.

  4. 04

    Gated pilot

    6–10 weeks

    What you keep

    A golden dataset labelled by your subject-matter experts, a regression suite wired into CI, live cost per successful task, and an acceptance threshold signed before the first result was seen.

    The decision it forces

    Proceed, iterate or kill — against a number written down and countersigned before the pilot started.

    When we stop

    Below the signed threshold after two iterations, we refund half the pilot-phase fee and you keep the golden dataset, the regression suite and the traces. Adoption is a separate gate: weekly active use under 30% of licensed seats, measured as the mean over the final four weeks of the pilot on a cohort of at least 15 seats — below that we switch it off, and the write-up goes to your sponsor and to ours with the reason named.

A gate that never stops anything is not a gate. It is decoration.

What you are probably thinking

We have already paid two consultancies for an AI strategy and nothing shipped.

Fair, and it is the norm rather than the exception — NANDA measured roughly 60% of tools investigated, 20% piloted, 5% implemented. The structural difference is what the engagement is contracted to produce: if the deliverable is a roadmap, you get a roadmap. Contract instead for a measured baseline, a portfolio with written kill criteria, and one use case through an acceptance gate whose threshold the process owner signed in advance. A useful test for any advisor is to ask for the last three use cases they recommended against building, and why. Here is our honest answer to our own test: Palamed has four delivered projects, not a hundred, so we cannot recite a decade of anonymised kills and we are not going to invent them. What you can hold us to is written into the engagement instead — a kill criterion and a decommission condition on every case before build, the rejected list handed over with the reason recorded, half the pilot-phase fee at risk against the signed threshold, and a week-five recommendation to build nothing if nothing clears twelve-month payback.

Our data is not ready. We have been told we need a data platform programme first.

This is how two years disappear. Enterprise-wide maturity is not a prerequisite for a specific use case; readiness of the critical data elements that case consumes is, and in practice that is often 5 to 15 fields with a named owner, documented lineage and an agreed quality SLA. We block the cases that genuinely cannot proceed and price the narrow remediation for the rest. Where the assessment shows the binding constraint is upstream master data rather than the model, the recommendation is to fix the fields and buy nothing — and that recommendation is on the table from week one, not after the invoice.

The AI Act was delayed, and Bulgaria has not even designated a market surveillance authority.

Both true, and neither is a reason to stop. The Omnibus moved standalone Annex III to 2 December 2027 and Annex I to 2 August 2028, but Article 5 prohibitions and Article 4 AI literacy have applied since 2 February 2025 and GPAI obligations since 2 August 2025. In Bulgaria the Ministry of Electronic Governance leads, Council of Ministers Decision No. 398 of 18 June 2025 designated seven fundamental-rights bodies, and as of August 2026 there is still no sanctions regime and no designated market surveillance authority. Obligations attach on the Regulation timetable regardless — and in practice your German and Nordic customers ask in a procurement questionnaire long before a regulator does.

Our engineers say they can build this internally.

Sometimes right, and the statistic usually quoted here is weaker than it looks. NANDA measured roughly 33% deployment success for internal builds against 67% for external partnerships, but that is survey data with an obvious confound: the organisations that hire outside help also have budget, an executive sponsor and a scoped problem, and those are the things that ship software. Build internally when three conditions are already true — someone owns the evaluation harness by name, the process owner can actually change the workflow, and there is a named maintenance owner for year two. Where any of the three is missing, buy the layer underneath and build only the thin differentiating part. Either way the decision should come from pricing build, buy and partner over 36 months with the same line items, human review and model migration included.

How do you prove ROI when the benefit is just people working faster?

By measuring before, not after — and this is the honest hard question, which is exactly why NANDA found budget flowing to sales and marketing, where attribution is easy, while back-office cases with faster payback went unfunded. A time study or a process-mining extract sets the baseline; a holdout group that does not get the tool separates your delta from seasonality. Value is then stated as cost per successful task against that baseline, and the ROI line counts only benefits a CFO will accept: reduced external spend, reduced overtime, or volume growth absorbed without hiring. If a case cannot be measured that way, we score it lower and say so.

We are regulated. One hallucination in a customer-facing answer and we have a problem.

Then the system should not be answering your customers unsupervised, and the design should say so on the first page. Most of what gets called hallucination in a RAG system is retrieval failure — the right passage was never in the context — which is measurable as recall@k and fixable. Beyond that: groundedness enforced as a release gate rather than a dashboard, citation-required output formats, confidence-based routing to human review, and a documented Article 14 human oversight arrangement naming a competent person rather than a role. The regulated deployment is a drafting or triage assistant under mandatory review, not an autonomous responder — and that version still carries most of the value, because the expensive part of the work was the reading, not the typing.

Why you rather than a Big Four firm?

Use whoever will put a falsifiable number in the contract. Ask any bidder for the article numbers governing your specific system, the post-Omnibus date each one applies from, and the acceptance threshold they would sign. Ask where the baseline comes from — process mining and a time study, or a manager estimate — and who builds the evaluation set. Those answers separate advisory from packaging at any firm size, and they are cheap for you to check. The difference here is smaller and more specific: the person who writes the strategy is the person who ships the first use case, and there is no pitch team handing you to juniors afterwards.

You are a small, founder-led firm. What happens if you are unavailable?

A fair question and the honest answer has limits in it. Palamed is founder-led, with named specialists brought in for the specific phase rather than a bench, and we cap this to two advisory engagements of this size at a time so the calendar is real. The protection is not headcount, it is the artifacts: every phase ends in something that reads without us — the Baseline Ledger, the AI register, the scored portfolio with the weights and the rejected list, the golden dataset and the regression suite in your repository, thresholds in writing. Any competent successor can pick that up. The limits we will not talk you out of: we do not staff a 24/7 on-call rota, and if you need forty consultants in a room next month we are the wrong call.

When we are the wrong choice

  • If you need a board-ready strategy deck this month, we are the wrong choice. Our first two weeks produce a measured baseline, which is slower, far less presentable, and the only thing that makes every later number falsifiable.
  • If no process owner has the authority to change how the work is actually done, an AI programme becomes a tooling purchase with a governance layer on top. We will say that at the first gate rather than build around an absent decision-maker.
  • If your largest manual process runs a few hundred transactions a month, the arithmetic rarely clears a twelve-month payback. The honest answer is usually a cleaner workflow or one fixed integration, and we would rather tell you that than sell a pilot.
  • Labour arbitrage is thin here. Bulgarian fully loaded labour runs around EUR 12.0 an hour against an EU average near EUR 34.9, which means a case that clears a twelve-month payback in Munich often will not in Sofia on the same volumes. We run that arithmetic in week two and tell you before you sign, not at the benefit review.

Questions we get asked

How do we choose between prompt engineering, RAG and fine-tuning?

In order of cost, not order of fashion. Start with retrieval when the answer already exists in your documents — and measure recall@k before touching the model, because chunking strategy is the parameter with the largest effect and the one almost nobody tunes. Fine-tune when you need a consistent output format or a domain register that prompting cannot hold, and only once a golden dataset exists to prove the delta. Fine-tuning on top of an unmeasured retrieval layer is an expensive way to be confidently wrong.

How do we tell whether one of our systems is high-risk?

Classify against Article 5 first, then Annex III and Annex I, and record the provider/deployer determination — deploying outside a vendor’s declared intended purpose, or substantially modifying a system, can turn a deployer into a provider. Article 6(3) offers four derogation conditions, but profiling of natural persons is always high-risk with no exception, and a system self-assessed as not high-risk under Article 6(3) must still be registered by its provider under Article 49(2). If the substantial modification is yours, that provider is you — and the filing is public and searchable, except law-enforcement, migration, asylum and border-control entries, which sit in the non-public section of the EU database under Article 49(4).

Can we use a frontier model without our data leaving the EU?

Usually, with a routed architecture: sensitive extraction and classification on an EU-hosted or self-hosted open-weight model, hard reasoning on an API with EU data residency and contractual no-training-on-inputs terms, and the sensitive fields never crossing the boundary. EDPB Opinion 28/2024 adds a step most vendors will not raise — a model trained on unlawfully processed personal data can face deployment restriction, so training-data provenance is a supplier due-diligence question, not only a hosting question.

What touchless rate should we plan for in our back office?

Document-heavy back offices typically start at 20–40% touchless. A credible first-year target is 60–80%; a vendor quoting 95% is usually redefining what counts as an exception. The number that decides whether headcount actually moves is the exception rate, and a large share of exceptions are upstream data errors — a wrong PO reference, missing master data — that no model fixes. We code an exception taxonomy from a real sample of several hundred items before anyone sizes the opportunity.

How much should we budget for evaluation and human review?

More than you expect, and it is the line that decides most cases. Human review commonly runs 5 to 50 times the inference cost, which makes cost control a workflow design problem rather than a procurement one. On the compute side, model tiering, semantic caching and output length limits commonly cut 30–60%. Continuous batching is already on if you are running vLLM or TGI, so the remaining compute levers in 2026 are prefix caching on shared system prompts, chunked prefill, FP8 weights and KV cache, and speculative decoding — typically 2–4× combined. We measure model FLOP utilisation rather than nvidia-smi utilisation, because the latter tells you a kernel is running, not that the GPU is doing useful work.

What do we do about staff already using personal AI accounts?

Treat it as discovery before you treat it as risk. NANDA found regular personal LLM use reported at over 90% of surveyed companies against roughly 40% holding an official subscription, so your AI inventory is already incomplete. Run a no-blame disclosure alongside SaaS and network log analysis, stand up a sanctioned tier with SSO, logging, EU residency and no-training terms, then fold the result into the AI Act register and the Article 4 literacy programme. The tasks people automate on their own are your best use-case signal.

Warm light ribbons on a dark field

What would you measure in our first two weeks?

Bring one process you suspect is expensive. We will tell you what we would time, which critical data elements it depends on, and whether it sits inside Annex III. If the answer is that the process itself is the problem, you will hear that first.

Scoped first step: the Baseline Ledger — two weeks, EUR 6,000–9,000 fixed, one process family. You keep the ledger and the AI register whether or not there is a phase two. End to end, baseline to a gated pilot in production, is 13–19 weeks; you can stop at any gate and keep what that phase produced.