Assessment of corporate AI sovereignty. Version 0.7.1 · August 2026 · CC BY 4.0
About this document
This field guide is the hands-on manual for AI-SAF-C, the business version of the AI Strategic Autonomy Framework, for a CTO’s office, a compliance team, or an AI director sizing up how self-reliant their company is and deciding what to fix first. For the concepts behind it — the two variants, capability × commitment, gap analysis, Version Deadlock — see Document 1; this guide assumes them. Dated market detail (which providers are certified, which chips ship) lives in Document 4, State of Play.
How each pillar is presented. Every pillar opens with what it measures and why, then a rubric table: sub-indicator, points, R/A tag, and the evidence expected. Score each sub-indicator as a whole number from 0 to its maximum; the pillar score is the sum. Report R and A subtotals with the headline (Framework §1.1).
Contents: Part I — the nine pillars. Part II — the ninth pillar in depth. Part III — scale discount. Part IV — worked example: BNP Paribas, with an ING calibration check. Part V — crosswalk to NIST AI RMF, ISO/IEC 42001, and GAIA-X.
Part I. The nine pillars
| # | Pillar | Weight | Corporate focus |
|---|---|---|---|
| 1 | Data | 14 | Own data, where it is stored, contractual and technical control of what leaves |
| 2 | Models | 16 | Portfolio across providers and jurisdictions; fine-tune capability; self-run fallback |
| 3 | Infrastructure | 14 | Cloud choice and jurisdiction, CLOUD Act exposure, own GPUs, confidential computing |
| 4 | Hardware | 10 | Reserved GPU capacity; vendor diversification; migration plan |
| 5 | Manufacturing | 1 | Supply-chain awareness only |
| 6 | Software | 10 | CUDA and tooling lock-in; open-source contribution |
| 7 | Human capital | 13 | AI-literate engineer share, retention, enablement programs |
| 8 | Regulatory | 12 | AI Act, DORA, GDPR, ISO 42001, model-risk management |
| 9 | Culture and operating model | 10 | Board oversight, applied literacy, speed, change governance, experimentation |
| Total | 100 |
Capable open-weights models have become cheap and widely available, so when everyone can get roughly the same models, your own private data — and the contracts and controls that protect it — becomes your strongest, hardest-to-copy advantage. The pillar scores four things: the corpus you hold, where it legally sits, what a provider may do with what you send, and what actually leaves your boundary at inference time.
No-train and data-rights discipline. The key contract terms: a clause barring the provider from using your inputs or outputs to train or tune models; retention limits (zero to 30 days); processing location (EU-only for sensitive data); named subprocessors; audit rights (at least annual SOC 2 Type II; on-site for regulated industries per DORA Article 30); ownership and deletion of fine-tune artifacts built from your data; and derived telemetry and embeddings under the same terms as the raw data.
Data-egress boundary control. Contracts govern what a provider may do; this governs whether data leaves at all — prompts, retrieved context (RAG), tool and agent calls, logs. IP and trade-secret exposure is scored separately: whether high-sensitivity IP (source code, designs, unreleased financials, M&A material) is kept out of, or isolated within, external model calls, with leakage monitoring. Synthetic data plays a limited role: fine for padding and test cases, never the core of fine-tuning data (Shumailov et al., Nature 2024 — models trained mostly on model output degrade irreversibly). A sovereign-inference fallback — what share of workloads could run in-house if egress had to stop tomorrow — is scored once, in the Models pillar, and credited here only by cross-reference.
| Sub-indicator | Points | R/A | Evidence expected |
|---|---|---|---|
| Own corpus and quality | 5 | A | Volume and purity of fine-tuning data; domain coverage |
| Residency and legal regime | 3 | R | Jurisdiction of storage and processing; CLOUD Act exposure; SCCs |
| Instruction and evaluation data | 1 | A | Own preference pairs; domain benchmarks |
| No-train and data-rights discipline | 2 | R | Clause quality; audit rights; retention; artifact ownership |
| Data-egress boundary control | 2 | R | Classification policy; AI gateway with logging; redaction/DLP; private inference for the most sensitive classes |
| IP and trade-secret exposure | 1 | R | Sensitive-IP classes blocked or isolated; leakage monitoring |
Pillar 2. Models — 16 points
Three questions: do you spread your bets across models, providers, and jurisdictions; can you fine-tune on your own data; and do you have a model you can run yourself if a Version Deadlock hits.
The five-level ladder, applied corporately (anchors L0–L4 = 0/4/11/13/16):
| Level | Points | Corporate profile | Note |
|---|---|---|---|
| L0 | 0 | Entirely on frontier APIs with no fallback | CLOUD Act + Version Deadlock exposure |
| L1 | 4 | On-prem open-weights fallback, no fine-tune | Physical control only |
| L2 | 11 | Product-level fine-tune on open weights | Fine-tune capability; open fallback |
| L3 | 13 | Domain-specific from-scratch models | Rare; only for very specialised needs |
| L4 | 16 | Foundation-model companies | Not themselves assessed on AI-SAF-C |
Most companies should aim for L2. The ladder anchors the level; points are earned through the sub-indicators, which sum to 16, with the anchor as reference.
Portfolio and jurisdiction diversification. The June 2026 Fable 5 / Mythos 5 suspension changed what diversification means. A portfolio of two or three providers all under the same jurisdiction is not diversified against the risk that actually materialised: a directive that in principle reached every customer of every US frontier vendor at once. The sub-indicator therefore scores legal regimes, not logos: US-jurisdiction frontier APIs, EU-jurisdiction providers, and self-hosted open weights (which carry no provider jurisdiction at inference) count as distinct classes. A portfolio of three US providers scores at most 2 of 5; provider count within a class adds at most 1 further point. A well-built portfolio runs three layers: a primary frontier service for the highest-quality work; a second provider — ideally in a different jurisdiction — shadow-tested quarterly; and a self-run open-weights model, fine-tuned to your domain, carrying real traffic daily rather than sitting idle. Real resilience degrades gradually rather than stopping dead.
| Sub-indicator | Points | R/A | Evidence expected |
|---|---|---|---|
| Portfolio and jurisdiction diversification | 5 | R | Distinct legal regimes in live use; active traffic split; ≤2 of 5 for single-jurisdiction portfolios |
| Fine-tune capability | 5 | A | L2 capability; fine-tuned models in production |
| Open-weights fallback | 4 | R | Actively operated on-prem/VPC instance; RTO < 72h; carries live traffic |
| Reasoning usage | 2 | A | Reasoning-class models integrated where they fit |
Pillar 3. Infrastructure and energy — 14 points
Four questions: which cloud and whose law; CLOUD Act exposure and mitigation; own GPU capacity; confidential-computing readiness.
The jurisdictional ladder for cloud choice runs, from strongest protection down: SecNumCloud 3.2 providers (immune to foreign-reach laws; required for French state clients); the CADA Cloud Sovereignty Framework levels as they are finalised (see Framework Part V); EU-operated subsidiaries of US providers (AWS European Sovereign Cloud — EU-run but still US-owned, so CLOUD Act exposure arguably remains); contractual EU boundaries (Azure EU Data Boundary — Microsoft’s own French Senate testimony conceded CLOUD Act protection cannot be guaranteed); and unmitigated US-jurisdiction cloud. Confidential computing (TEEs: Intel TDX, AMD SEV-SNP, NVIDIA Confidential Computing) lets a regulated EU company run on a US cloud with weights and inputs encrypted in use — credited as genuine mitigation. Current certified providers and capacity figures: Document 4.
| Sub-indicator | Points | R/A | Evidence expected |
|---|---|---|---|
| Cloud choice and jurisdiction | 5 | R | Certified-sovereign, EU-subsidiary, or documented mitigation strategy |
| On-prem / in-VPC GPU | 4 | R | Own or reserved capacity in H100-equivalents |
| Network and backup independence | 2 | R | Multi-region, multi-vendor; tested DR runbooks |
| Confidential-computing readiness | 3 | R | TEE patterns on regulated workloads; attestation; HSM integration |
Pillar 4. Hardware architecture — 10 points
Three questions: how much guaranteed GPU capacity is locked in for a crunch; how reliant on one supplier; is there a tested plan to switch. NVIDIA still dominates, but AMD is a believable second source and cloud ASICs (TPU, Trainium) a third class; details in Document 4.
| Sub-indicator | Points | R/A | Evidence expected |
|---|---|---|---|
| Reserved GPU capacity | 4 | R | Multi-year reservations in H100-equivalents; T+3 projection |
| Vendor diversification | 4 | R | Production traffic on at least two accelerator vendors |
| Migration plan under Version Deadlock | 2 | R | Documented, tested fallback under export-control scenarios |
Pillar 5. Manufacturing — 1 point
A company that does not build foundation models cannot meaningfully score here. One point [R] is given for demonstrated supply-chain awareness — knowing where the packaging and memory bottlenecks are and feeding that into buying decisions. The single point is kept as a separate pillar so the eight shared pillars stay one-to-one with the national variant.
Pillar 6. Software stack — 10 points
Worth two points more than in the national variant because companies feel lock-in in everyday work: CUDA exposure, AI-ops tooling, and inference-runtime choice each directly affect Version Deadlock risk. The runtime landscape and CUDA-alternative status: Document 4.
| Sub-indicator | Points | R/A | Evidence expected |
|---|---|---|---|
| CUDA exposure | 3 | R | Share of production workloads dependent on CUDA alone |
| Training-framework diversification | 2 | R | PyTorch + JAX or alternatives; reproducibility discipline |
| Inference-runtime diversification | 2 | R | vLLM / SGLang / llama.cpp alongside or instead of TensorRT |
| MLOps lock-in | 2 | R | Vendor-agnostic deployment; portability demonstrated |
| Open-source contribution | 1 | A | Upstream commits, not merely consumption |
Pillar 7. Human capital — 13 points
The pillar with the biggest knock-on effect: gains here improve the payoff from every other investment. BCG’s survey work suggests employee training is the single best predictor of AI-program success (strong-results share rising from 12% to 28%).
| Sub-indicator | Points | R/A | Evidence expected |
|---|---|---|---|
| AI-literate engineer ratio | 4 | A | Share of engineers actively working with AI tooling |
| ML/AI specialists | 3 | A | Senior ML engineers and researchers on staff |
| AI Champions / enablement program | 3 | A | Formal cross-functional program (the AI Factory pattern) |
| Retention against brain drain | 2 | R | AI-staff retention metrics; competitive compensation |
| AI literacy in non-tech staff | 1 | A | Formal training coverage; internal-assistant adoption |
Pillar 8. Regulatory (compliance) — 12 points
For businesses this pillar measures how thoroughly you follow the rules — the mirror image of the national variant (Framework §1.2). The reference obligations: the EU AI Act (GPAI obligations and Article 50 transparency enforceable since 2 August 2026; high-risk obligations deferred by the Digital Omnibus to December 2027 for Annex III systems and August 2028 for Annex I products — deadlines that feed the commitment projection, not an excuse to stop preparing); DORA (January 2025) for financial firms; GDPR plus Data Act Chapter VII (September 2025) for residency and CLOUD Act shielding; SR 11-7-style model-risk management; and ISO/IEC 42001 certification, credited as proof of disciplined procedure.
| Sub-indicator | Points | R/A | Evidence expected |
|---|---|---|---|
| AI Act conformity posture | 4 | A | Risk classification complete; technical documentation; FRIA; registration |
| DPIA + GDPR + Data Act compliance | 2 | A | DPIA completeness; lawful basis; SCCs; Chapter VII compatibility |
| MRM (SR 11-7 equivalent) | 3 | R | Model inventory; challenger models; monitoring; independent validation |
| ISO 42001 certification | 2 | A | Active certification; SoA completeness; audit evidence |
| DORA readiness (financial) | 1 | R | ICT risk management; third-party register; incident reporting |
Part II. The ninth pillar — Culture and Operating Model
Getting AI right in a company is not just technology and rule-following. The CrowdStrike incident (July 2024) was the clearest recent test of operational maturity: whether recovery took hours or days came down to playbooks and response teams, not cloud choice. Every major business AI-maturity framework (McKinsey, BCG, Gartner, Deloitte, MIT CISR) treats culture as its own dimension, and for the same reason AI-SAF-C gives it a 10-point pillar. “Culture” here means how the company actually operates — measured by what people visibly do (real AI use in everyday work, deployment speed, escalation, behaviour change after incidents), not by mission statements or survey answers.
The ten points split evenly — five for governance, five for practice — on purpose. Governance without practice produces polished policies that people route around; practice without governance produces companies that ship fast and discover accountability only when something breaks.
Governance side (5):
| Sub-indicator | Points | R/A | Scoring anchors |
|---|---|---|---|
| AI ethics / model-risk committee | 3 | A | 0 = none → 3 = board-level chair; monthly or on-demand; tied into risk and audit; full risk-ranked model inventory |
| Written model-risk management | 2 | A | 0 = no policy → 2 = documented checks per live model, monitoring, retirement schedule, staffed independent validation |
Practice side (5):
| Sub-indicator | Points | R/A | Scoring anchors |
|---|---|---|---|
| Incident handling | 2 | R | 0 = no AI-specific process → 2 = rehearsed response with drills, tested rollback, routine near-miss reporting |
| Rollout speed | 1 | R | 0 = >6 months or untracked → 1 = ≤3 months, tracked against a target |
| Disciplined experimenting | 2 | A | 0 = no systematic testing → 2 = own-domain benchmarks; A/B tooling wired into rollout; results gate go-live |
Calibration: a long-established, heavily regulated firm with strong governance but regulation-limited speed scores around 8; a mature AI-first scaleup with strong practice but lighter governance around 7; an early-stage firm with neither, 1–2. The even split keeps either type from winning the pillar on company type alone.
Part III. Scale discount
Size affects how realistically some scores can be reached. For enterprises under 500 employees, the Data, Human Capital, and Culture raw scores are multiplied by 0.85; from 500 to 2,000 employees, by 0.92; above 2,000, no discount. The other six pillars are not adjusted — they depend on architectural and contractual choices, not headcount. A 30-person startup cannot build a bank-sized fine-tuning corpus however well it manages data, and formal governance comes more easily at scale; but a small firm can reach L2 on Models, run on certified sovereign infrastructure, hold ISO 42001, and diversify hardware — none of that needs a thousand people.
Part IV. Worked example: BNP Paribas
The strongest example of a European bank doing AI well: a multi-year Mistral partnership since February 2024; an in-house “LLM as a service” platform on BNP’s own GPUs in its own data centres; disciplined governance. (Both profiles below are the author’s estimates from public information — see the Evidence Base’s verification note.)
4.1 Base profile
| Pillar | Weight | Raw | Score | Note |
|---|---|---|---|---|
| Data | 14 | 0.70 | 9.8 | Large proprietary financial datasets; strict no-train contracts |
| Models | 16 | 0.55 | 8.8 | L2: Mistral partnership + internal fine-tunes; genuine jurisdiction spread (EU provider + self-run weights); no own FM |
| Infrastructure | 14 | 0.70 | 9.8 | LLM-aaS on own data centres; backer of Mistral’s Essonne DC |
| Hardware | 10 | 0.35 | 3.5 | Some H100/H200 in-house; NVIDIA-only |
| Manufacturing | 1 | 0.00 | 0.0 | Not applicable |
| Software | 10 | 0.55 | 5.5 | Mature MLOps; open weights + proprietary mix |
| Human capital | 13 | 0.65 | 8.5 | AI Factory / AI Champion programs; large quant population |
| Regulatory | 12 | 0.85 | 10.2 | DORA, AI Act readiness, GDPR-native, SR 11-7-equivalent MRM |
| Culture | 10 | 0.55 | 5.5 | Board-level oversight; speed limited by risk committees |
| Total | 100 | ~61.6 | No scale discount (>2,000 employees) |
4.2 Reading the profile
Regulatory is the strongest pillar — banking is the most heavily regulated industry there is, and BNP has two centuries of practice at it. Data is second: huge private financial datasets that no non-bank could assemble, protected by tight no-train contracts. Infrastructure scores on the on-prem platform and the unusual depth of the Mistral relationship — BNP is at once customer and financier. Models has growth room: L2 is the right level for a bank (building a foundation model is not a bank’s job), and BNP’s portfolio is one of the few in EU banking with real jurisdictional spread — an EU frontier provider plus self-run open weights — which is what the diversification sub-indicator now rewards. Hardware and Culture are the constraints: NVIDIA dependence is industry-wide, and a big regulated bank deploys slowly by design.
4.3 Calibration comparison: ING Bank
ING made a different choice: primarily Azure OpenAI, less in its own data centres, more CLOUD Act exposure — and, in the revised Models terms, a single-jurisdiction portfolio.
| Pillar | Weight | Raw | Score | Note |
|---|---|---|---|---|
| Data | 14 | 0.55 | 7.7 | Strong datasets; contractual discipline limited |
| Models | 16 | 0.35 | 5.6 | Primarily Azure OpenAI; single jurisdiction; limited fine-tune; insufficient fallback |
| Infrastructure | 14 | 0.40 | 5.6 | Azure EU Data Boundary; CLOUD Act exposure; limited on-prem |
| Hardware | 10 | 0.20 | 2.0 | Minimal in-house; fully cloud dependent |
| Manufacturing | 1 | 0.00 | 0.0 | N/A |
| Software | 10 | 0.50 | 5.0 | Standard enterprise stack |
| Human capital | 13 | 0.55 | 7.2 | Smaller AI staff than BNP |
| Regulatory | 12 | 0.75 | 9.0 | DORA native; AI Act readiness in progress |
| Culture | 10 | 0.50 | 5.0 | Moderately slower pace |
| Total | 100 | ~47.1 |
4.4 A structural finding for EU banking
Every EU bank scores at least 35 on compliance depth alone — DORA, the AI Act, GDPR, and MRM; a bank below 35 will draw supervisory action. For mid-sized EU banks the practical range is 45–60. Above 60 requires either a deep partnership with an EU model company or a large self-run model estate — today BNP, potentially Santander, HSBC, or UBS with the same investment. Above 70 would require building a top-tier model from scratch, which is not a realistic goal for a bank.
Part V. Crosswalk to business standards
AI-SAF-C is not a separate audit. It is a higher-level score that reuses evidence you already have — NIST AI RMF playbooks, ISO/IEC 42001 Statements of Applicability, GAIA-X self-descriptions — turning “yet another assessment” into a reuse of compliance work already paid for.
| AI-SAF-C pillar | NIST AI RMF | ISO 42001 Annex A | GAIA-X criterion |
|---|---|---|---|
| Data (14) | MAP 2.3, MAP 3, MEASURE 2.8/2.10 | A.7 Data + A.4 | Data protection; Portability; European control |
| Models (16) | MEASURE (all) + MANAGE 4.1; GenAI Profile | A.6 Lifecycle + A.5 | Transparency |
| Infrastructure (14) | GOVERN 6, MAP 4, MANAGE 3 | A.4 Resources + A.10 Third-party | Security + European control (L2/L3) |
| Hardware (10) | GOVERN 6.1 | A.4 + A.10 | European control |
| Manufacturing (1) | — | — | — |
| Software (10) | GOVERN 1.5, MEASURE 2.7, MANAGE 2.3 | A.6 + A.10 | Security (SDL controls) |
| Human capital (13) | GOVERN 2, 3, 4 | A.3 + A.4.5 | Transparency (partial) |
| Regulatory (12) | GOVERN 1 + MAP 1 | A.2 + A.5 + A.8 | Contractual governance + Data protection |
| Culture (10) | GOVERN 3 | A.3 Internal organisation | — |
How to use it: gather the compliance paperwork you already hold; pull the matching evidence per pillar via the table; answer the few questions the paperwork does not cover (mainly: a workable self-run fallback, no-train discipline, deployment and experimentation speed); report the AI-SAF-C score alongside the existing compliance reports, not as a separate document. For EU companies serving regulated customers, GAIA-X Level 3 feeds straight into the Infrastructure score; Level 2 is a sensible minimum for business software sold to non-regulated customers. As the CADA Cloud Sovereignty Framework is finalised, its assurance levels will slot into the same column.
For the master glossary, see Document 1. Terms used here without definition are defined there.