AI-SAF-C Field Guide — Business variant

Assessment of corporate AI sovereignty. Version 0.7.1 · August 2026 · CC BY 4.0

About this document

This field guide is the hands-on manual for AI-SAF-C, the business version of the AI Strategic Autonomy Framework, for a CTO’s office, a compliance team, or an AI director sizing up how self-reliant their company is and deciding what to fix first. For the concepts behind it — the two variants, capability × commitment, gap analysis, Version Deadlock — see Document 1; this guide assumes them. Dated market detail (which providers are certified, which chips ship) lives in Document 4, State of Play.

How each pillar is presented. Every pillar opens with what it measures and why, then a rubric table: sub-indicator, points, R/A tag, and the evidence expected. Score each sub-indicator as a whole number from 0 to its maximum; the pillar score is the sum. Report R and A subtotals with the headline (Framework §1.1).

Contents: Part I — the nine pillars. Part II — the ninth pillar in depth. Part III — scale discount. Part IV — worked example: BNP Paribas, with an ING calibration check. Part V — crosswalk to NIST AI RMF, ISO/IEC 42001, and GAIA-X.

Part I. The nine pillars

#PillarWeightCorporate focus
1Data14Own data, where it is stored, contractual and technical control of what leaves
2Models16Portfolio across providers and jurisdictions; fine-tune capability; self-run fallback
3Infrastructure14Cloud choice and jurisdiction, CLOUD Act exposure, own GPUs, confidential computing
4Hardware10Reserved GPU capacity; vendor diversification; migration plan
5Manufacturing1Supply-chain awareness only
6Software10CUDA and tooling lock-in; open-source contribution
7Human capital13AI-literate engineer share, retention, enablement programs
8Regulatory12AI Act, DORA, GDPR, ISO 42001, model-risk management
9Culture and operating model10Board oversight, applied literacy, speed, change governance, experimentation
Total100

Capable open-weights models have become cheap and widely available, so when everyone can get roughly the same models, your own private data — and the contracts and controls that protect it — becomes your strongest, hardest-to-copy advantage. The pillar scores four things: the corpus you hold, where it legally sits, what a provider may do with what you send, and what actually leaves your boundary at inference time.

No-train and data-rights discipline. The key contract terms: a clause barring the provider from using your inputs or outputs to train or tune models; retention limits (zero to 30 days); processing location (EU-only for sensitive data); named subprocessors; audit rights (at least annual SOC 2 Type II; on-site for regulated industries per DORA Article 30); ownership and deletion of fine-tune artifacts built from your data; and derived telemetry and embeddings under the same terms as the raw data.

Data-egress boundary control. Contracts govern what a provider may do; this governs whether data leaves at all — prompts, retrieved context (RAG), tool and agent calls, logs. IP and trade-secret exposure is scored separately: whether high-sensitivity IP (source code, designs, unreleased financials, M&A material) is kept out of, or isolated within, external model calls, with leakage monitoring. Synthetic data plays a limited role: fine for padding and test cases, never the core of fine-tuning data (Shumailov et al., Nature 2024 — models trained mostly on model output degrade irreversibly). A sovereign-inference fallback — what share of workloads could run in-house if egress had to stop tomorrow — is scored once, in the Models pillar, and credited here only by cross-reference.

Sub-indicatorPointsR/AEvidence expected
Own corpus and quality5AVolume and purity of fine-tuning data; domain coverage
Residency and legal regime3RJurisdiction of storage and processing; CLOUD Act exposure; SCCs
Instruction and evaluation data1AOwn preference pairs; domain benchmarks
No-train and data-rights discipline2RClause quality; audit rights; retention; artifact ownership
Data-egress boundary control2RClassification policy; AI gateway with logging; redaction/DLP; private inference for the most sensitive classes
IP and trade-secret exposure1RSensitive-IP classes blocked or isolated; leakage monitoring

Pillar 2. Models — 16 points

Three questions: do you spread your bets across models, providers, and jurisdictions; can you fine-tune on your own data; and do you have a model you can run yourself if a Version Deadlock hits.

The five-level ladder, applied corporately (anchors L0–L4 = 0/4/11/13/16):

LevelPointsCorporate profileNote
L00Entirely on frontier APIs with no fallbackCLOUD Act + Version Deadlock exposure
L14On-prem open-weights fallback, no fine-tunePhysical control only
L211Product-level fine-tune on open weightsFine-tune capability; open fallback
L313Domain-specific from-scratch modelsRare; only for very specialised needs
L416Foundation-model companiesNot themselves assessed on AI-SAF-C

Most companies should aim for L2. The ladder anchors the level; points are earned through the sub-indicators, which sum to 16, with the anchor as reference.

Portfolio and jurisdiction diversification. The June 2026 Fable 5 / Mythos 5 suspension changed what diversification means. A portfolio of two or three providers all under the same jurisdiction is not diversified against the risk that actually materialised: a directive that in principle reached every customer of every US frontier vendor at once. The sub-indicator therefore scores legal regimes, not logos: US-jurisdiction frontier APIs, EU-jurisdiction providers, and self-hosted open weights (which carry no provider jurisdiction at inference) count as distinct classes. A portfolio of three US providers scores at most 2 of 5; provider count within a class adds at most 1 further point. A well-built portfolio runs three layers: a primary frontier service for the highest-quality work; a second provider — ideally in a different jurisdiction — shadow-tested quarterly; and a self-run open-weights model, fine-tuned to your domain, carrying real traffic daily rather than sitting idle. Real resilience degrades gradually rather than stopping dead.

Sub-indicatorPointsR/AEvidence expected
Portfolio and jurisdiction diversification5RDistinct legal regimes in live use; active traffic split; ≤2 of 5 for single-jurisdiction portfolios
Fine-tune capability5AL2 capability; fine-tuned models in production
Open-weights fallback4RActively operated on-prem/VPC instance; RTO < 72h; carries live traffic
Reasoning usage2AReasoning-class models integrated where they fit

Pillar 3. Infrastructure and energy — 14 points

Four questions: which cloud and whose law; CLOUD Act exposure and mitigation; own GPU capacity; confidential-computing readiness.

The jurisdictional ladder for cloud choice runs, from strongest protection down: SecNumCloud 3.2 providers (immune to foreign-reach laws; required for French state clients); the CADA Cloud Sovereignty Framework levels as they are finalised (see Framework Part V); EU-operated subsidiaries of US providers (AWS European Sovereign Cloud — EU-run but still US-owned, so CLOUD Act exposure arguably remains); contractual EU boundaries (Azure EU Data Boundary — Microsoft’s own French Senate testimony conceded CLOUD Act protection cannot be guaranteed); and unmitigated US-jurisdiction cloud. Confidential computing (TEEs: Intel TDX, AMD SEV-SNP, NVIDIA Confidential Computing) lets a regulated EU company run on a US cloud with weights and inputs encrypted in use — credited as genuine mitigation. Current certified providers and capacity figures: Document 4.

Sub-indicatorPointsR/AEvidence expected
Cloud choice and jurisdiction5RCertified-sovereign, EU-subsidiary, or documented mitigation strategy
On-prem / in-VPC GPU4ROwn or reserved capacity in H100-equivalents
Network and backup independence2RMulti-region, multi-vendor; tested DR runbooks
Confidential-computing readiness3RTEE patterns on regulated workloads; attestation; HSM integration

Pillar 4. Hardware architecture — 10 points

Three questions: how much guaranteed GPU capacity is locked in for a crunch; how reliant on one supplier; is there a tested plan to switch. NVIDIA still dominates, but AMD is a believable second source and cloud ASICs (TPU, Trainium) a third class; details in Document 4.

Sub-indicatorPointsR/AEvidence expected
Reserved GPU capacity4RMulti-year reservations in H100-equivalents; T+3 projection
Vendor diversification4RProduction traffic on at least two accelerator vendors
Migration plan under Version Deadlock2RDocumented, tested fallback under export-control scenarios

Pillar 5. Manufacturing — 1 point

A company that does not build foundation models cannot meaningfully score here. One point [R] is given for demonstrated supply-chain awareness — knowing where the packaging and memory bottlenecks are and feeding that into buying decisions. The single point is kept as a separate pillar so the eight shared pillars stay one-to-one with the national variant.

Pillar 6. Software stack — 10 points

Worth two points more than in the national variant because companies feel lock-in in everyday work: CUDA exposure, AI-ops tooling, and inference-runtime choice each directly affect Version Deadlock risk. The runtime landscape and CUDA-alternative status: Document 4.

Sub-indicatorPointsR/AEvidence expected
CUDA exposure3RShare of production workloads dependent on CUDA alone
Training-framework diversification2RPyTorch + JAX or alternatives; reproducibility discipline
Inference-runtime diversification2RvLLM / SGLang / llama.cpp alongside or instead of TensorRT
MLOps lock-in2RVendor-agnostic deployment; portability demonstrated
Open-source contribution1AUpstream commits, not merely consumption

Pillar 7. Human capital — 13 points

The pillar with the biggest knock-on effect: gains here improve the payoff from every other investment. BCG’s survey work suggests employee training is the single best predictor of AI-program success (strong-results share rising from 12% to 28%).

Sub-indicatorPointsR/AEvidence expected
AI-literate engineer ratio4AShare of engineers actively working with AI tooling
ML/AI specialists3ASenior ML engineers and researchers on staff
AI Champions / enablement program3AFormal cross-functional program (the AI Factory pattern)
Retention against brain drain2RAI-staff retention metrics; competitive compensation
AI literacy in non-tech staff1AFormal training coverage; internal-assistant adoption

Pillar 8. Regulatory (compliance) — 12 points

For businesses this pillar measures how thoroughly you follow the rules — the mirror image of the national variant (Framework §1.2). The reference obligations: the EU AI Act (GPAI obligations and Article 50 transparency enforceable since 2 August 2026; high-risk obligations deferred by the Digital Omnibus to December 2027 for Annex III systems and August 2028 for Annex I products — deadlines that feed the commitment projection, not an excuse to stop preparing); DORA (January 2025) for financial firms; GDPR plus Data Act Chapter VII (September 2025) for residency and CLOUD Act shielding; SR 11-7-style model-risk management; and ISO/IEC 42001 certification, credited as proof of disciplined procedure.

Sub-indicatorPointsR/AEvidence expected
AI Act conformity posture4ARisk classification complete; technical documentation; FRIA; registration
DPIA + GDPR + Data Act compliance2ADPIA completeness; lawful basis; SCCs; Chapter VII compatibility
MRM (SR 11-7 equivalent)3RModel inventory; challenger models; monitoring; independent validation
ISO 42001 certification2AActive certification; SoA completeness; audit evidence
DORA readiness (financial)1RICT risk management; third-party register; incident reporting

Part II. The ninth pillar — Culture and Operating Model

Getting AI right in a company is not just technology and rule-following. The CrowdStrike incident (July 2024) was the clearest recent test of operational maturity: whether recovery took hours or days came down to playbooks and response teams, not cloud choice. Every major business AI-maturity framework (McKinsey, BCG, Gartner, Deloitte, MIT CISR) treats culture as its own dimension, and for the same reason AI-SAF-C gives it a 10-point pillar. “Culture” here means how the company actually operates — measured by what people visibly do (real AI use in everyday work, deployment speed, escalation, behaviour change after incidents), not by mission statements or survey answers.

The ten points split evenly — five for governance, five for practice — on purpose. Governance without practice produces polished policies that people route around; practice without governance produces companies that ship fast and discover accountability only when something breaks.

Governance side (5):

Sub-indicatorPointsR/AScoring anchors
AI ethics / model-risk committee3A0 = none → 3 = board-level chair; monthly or on-demand; tied into risk and audit; full risk-ranked model inventory
Written model-risk management2A0 = no policy → 2 = documented checks per live model, monitoring, retirement schedule, staffed independent validation

Practice side (5):

Sub-indicatorPointsR/AScoring anchors
Incident handling2R0 = no AI-specific process → 2 = rehearsed response with drills, tested rollback, routine near-miss reporting
Rollout speed1R0 = >6 months or untracked → 1 = ≤3 months, tracked against a target
Disciplined experimenting2A0 = no systematic testing → 2 = own-domain benchmarks; A/B tooling wired into rollout; results gate go-live

Calibration: a long-established, heavily regulated firm with strong governance but regulation-limited speed scores around 8; a mature AI-first scaleup with strong practice but lighter governance around 7; an early-stage firm with neither, 1–2. The even split keeps either type from winning the pillar on company type alone.

Part III. Scale discount

Size affects how realistically some scores can be reached. For enterprises under 500 employees, the Data, Human Capital, and Culture raw scores are multiplied by 0.85; from 500 to 2,000 employees, by 0.92; above 2,000, no discount. The other six pillars are not adjusted — they depend on architectural and contractual choices, not headcount. A 30-person startup cannot build a bank-sized fine-tuning corpus however well it manages data, and formal governance comes more easily at scale; but a small firm can reach L2 on Models, run on certified sovereign infrastructure, hold ISO 42001, and diversify hardware — none of that needs a thousand people.

Part IV. Worked example: BNP Paribas

The strongest example of a European bank doing AI well: a multi-year Mistral partnership since February 2024; an in-house “LLM as a service” platform on BNP’s own GPUs in its own data centres; disciplined governance. (Both profiles below are the author’s estimates from public information — see the Evidence Base’s verification note.)

4.1 Base profile

PillarWeightRawScoreNote
Data140.709.8Large proprietary financial datasets; strict no-train contracts
Models160.558.8L2: Mistral partnership + internal fine-tunes; genuine jurisdiction spread (EU provider + self-run weights); no own FM
Infrastructure140.709.8LLM-aaS on own data centres; backer of Mistral’s Essonne DC
Hardware100.353.5Some H100/H200 in-house; NVIDIA-only
Manufacturing10.000.0Not applicable
Software100.555.5Mature MLOps; open weights + proprietary mix
Human capital130.658.5AI Factory / AI Champion programs; large quant population
Regulatory120.8510.2DORA, AI Act readiness, GDPR-native, SR 11-7-equivalent MRM
Culture100.555.5Board-level oversight; speed limited by risk committees
Total100~61.6No scale discount (>2,000 employees)

4.2 Reading the profile

Regulatory is the strongest pillar — banking is the most heavily regulated industry there is, and BNP has two centuries of practice at it. Data is second: huge private financial datasets that no non-bank could assemble, protected by tight no-train contracts. Infrastructure scores on the on-prem platform and the unusual depth of the Mistral relationship — BNP is at once customer and financier. Models has growth room: L2 is the right level for a bank (building a foundation model is not a bank’s job), and BNP’s portfolio is one of the few in EU banking with real jurisdictional spread — an EU frontier provider plus self-run open weights — which is what the diversification sub-indicator now rewards. Hardware and Culture are the constraints: NVIDIA dependence is industry-wide, and a big regulated bank deploys slowly by design.

4.3 Calibration comparison: ING Bank

ING made a different choice: primarily Azure OpenAI, less in its own data centres, more CLOUD Act exposure — and, in the revised Models terms, a single-jurisdiction portfolio.

PillarWeightRawScoreNote
Data140.557.7Strong datasets; contractual discipline limited
Models160.355.6Primarily Azure OpenAI; single jurisdiction; limited fine-tune; insufficient fallback
Infrastructure140.405.6Azure EU Data Boundary; CLOUD Act exposure; limited on-prem
Hardware100.202.0Minimal in-house; fully cloud dependent
Manufacturing10.000.0N/A
Software100.505.0Standard enterprise stack
Human capital130.557.2Smaller AI staff than BNP
Regulatory120.759.0DORA native; AI Act readiness in progress
Culture100.505.0Moderately slower pace
Total100~47.1

4.4 A structural finding for EU banking

Every EU bank scores at least 35 on compliance depth alone — DORA, the AI Act, GDPR, and MRM; a bank below 35 will draw supervisory action. For mid-sized EU banks the practical range is 45–60. Above 60 requires either a deep partnership with an EU model company or a large self-run model estate — today BNP, potentially Santander, HSBC, or UBS with the same investment. Above 70 would require building a top-tier model from scratch, which is not a realistic goal for a bank.

Part V. Crosswalk to business standards

AI-SAF-C is not a separate audit. It is a higher-level score that reuses evidence you already have — NIST AI RMF playbooks, ISO/IEC 42001 Statements of Applicability, GAIA-X self-descriptions — turning “yet another assessment” into a reuse of compliance work already paid for.

AI-SAF-C pillarNIST AI RMFISO 42001 Annex AGAIA-X criterion
Data (14)MAP 2.3, MAP 3, MEASURE 2.8/2.10A.7 Data + A.4Data protection; Portability; European control
Models (16)MEASURE (all) + MANAGE 4.1; GenAI ProfileA.6 Lifecycle + A.5Transparency
Infrastructure (14)GOVERN 6, MAP 4, MANAGE 3A.4 Resources + A.10 Third-partySecurity + European control (L2/L3)
Hardware (10)GOVERN 6.1A.4 + A.10European control
Manufacturing (1)———
Software (10)GOVERN 1.5, MEASURE 2.7, MANAGE 2.3A.6 + A.10Security (SDL controls)
Human capital (13)GOVERN 2, 3, 4A.3 + A.4.5Transparency (partial)
Regulatory (12)GOVERN 1 + MAP 1A.2 + A.5 + A.8Contractual governance + Data protection
Culture (10)GOVERN 3A.3 Internal organisation—

How to use it: gather the compliance paperwork you already hold; pull the matching evidence per pillar via the table; answer the few questions the paperwork does not cover (mainly: a workable self-run fallback, no-train discipline, deployment and experimentation speed); report the AI-SAF-C score alongside the existing compliance reports, not as a separate document. For EU companies serving regulated customers, GAIA-X Level 3 feeds straight into the Infrastructure score; Level 2 is a sensible minimum for business software sold to non-regulated customers. As the CADA Cloud Sovereignty Framework is finalised, its assurance levels will slot into the same column.

For the master glossary, see Document 1. Terms used here without definition are defined there.