• Towards AGI
  • Posts
  • The Pentagon Unplugged Claude. Could Your AI Survive A Vendor Switch?

The Pentagon Unplugged Claude. Could Your AI Survive A Vendor Switch?

why owning your context matters more than picking your model

Today, we're diving into:

  • Hot Tea: The Pentagon has finally unplugged Claude

  • Gen AI: Europe ships a trillion-parameter model you can run yourself

  • OpenAI: 372 new math results, and a verification problem

The Pentagon Unplugged Claude. It Took Months Longer Than Planned.

Switching AI vendors sounds like a procurement decision. The Pentagon just showed it is closer to an organ transplant.

On October 5, the Department of Defense said it has ceased the use of Anthropic products, months after Defense Secretary Pete Hegseth designated the company a supply chain risk and set a six-month phase-out. The late-August deadline came and went. Sources said Claude was still in use the week before the announcement.

Could You Switch Models In Six Months?

Find out how tightly your AI is wired to one vendor before a contract, a regulator or a price change forces the question.

An enterprise swapping one AI model for another at its central data hub

Why The Exit Took So Long

The dispute started in February, when Anthropic refused to remove safety guardrails from Claude. The phase-out order followed. Then reality set in.

Claude was embedded in Maven Smart System, the Palantir-run platform the Pentagon uses as its primary intelligence data system. You do not lift a model out of a workflow like that by changing an API key.

"Once they become integrated it can be painful to remove them." Lauren Kahn, Georgetown's Center for Security and Emerging Technology

Kahn's point was simple: these systems are not plug and play. The Pentagon has since signed with Google, xAI and OpenAI. It now has to rebuild what Claude was doing inside each of those workflows.

This Is Not A Defense Story

Your organization will probably never be blacklisted by a government. But every enterprise faces the same forces at a smaller scale: a vendor changes its pricing, a regulator questions a provider, a model gets retired, or a better option appears.

The real bottleneck is not the model. It is everything you built around it.

Where The Lock-In Actually Lives

Most leaders think about lock-in at the contract level. The deeper lock-in sits in places nobody inventories.

  • Prompts and agent logic tuned to one model's quirks.

  • Business context stored inside a vendor's memory, workspace or vector store.

  • Evaluation baselines that only exist for the model you use today.

  • Partner platforms that call a model you never chose directly.

The wrong reaction is to sign with every vendor and hope that counts as a strategy. Three contracts with three providers just gives you three places to be locked in.

The Better Alternative

Own the context. Rent the model.

Look at what Atlassian announced the next day. Its expanded OpenAI partnership puts GPT-6 Astra, the GPT-5.6 series and Codex behind its Rovo agents. But the intelligence runs on top of Atlassian's Teamwork Graph, which it describes as an enterprise context layer connecting people, projects, documents and decisions.

The models are powerful. The context layer is Atlassian's own. That is the part that does not have to move when a model does.

Building that kind of portability starts with trusted enterprise context. DataManagement.AI helps organizations create that foundation by unifying metadata, lineage, governance, business glossaries and knowledge assets into a single intelligence layer, so your context stays yours and any approved model can plug into it.

A governed data foundation with interchangeable AI models on top

What You Should Do This Quarter

  • Map every workflow that depends on one model. Include the ones inside partner platforms.

  • Keep business context outside any vendor. Your glossary, lineage and policies should not live in someone else's memory feature.

  • Build a model-neutral evaluation set. You cannot switch safely if you cannot compare.

  • Write an exit clause and a timeline. Then test it once, before you need it.

The organizations that win the model race will not be the ones that pick the right vendor. They will be the ones that can change vendors in weeks, not seasons.

Europe Just Shipped A Trillion-Parameter Model You Can Run Yourself

For two years, the frontier conversation has been a two-country story. Mistral just made a serious case for a third option.

On October 6, the Paris-based lab introduced Mistral Large 4, nicknamed Le Chonk: a natively multimodal model with 1 trillion parameters, 49 billion of them active on any given token. Mistral says it was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in its own datacenters in Europe.

A sovereign AI model running inside a protected data center

Sovereignty Just Became A Procurement Option

The headline is not the parameter count. It is the deployment model.

Mistral says the model is "designed to give customers control over their AI," and can be deployed in Europe "independently of other digital service providers and under European law."

The preview API is live now. Open weights are promised by the end of October, which means you will be able to run a frontier-class model inside your own environment, under your own policies.

Who Should Pay Attention

If you operate in a regulated industry, handle data that cannot leave a jurisdiction, or sell to European public sector buyers, this changes your options. Until now, the strongest models came with a simple trade: capability in exchange for sending your data to someone else's cloud.

That trade is getting weaker. Mistral says it built the model with enterprises in finance, manufacturing, pharmaceuticals, logistics and the public sector, and trained it on more than 160 languages.

Do Not Mistake Open For Free

Open weights remove the license fee. They do not remove the bill.

A trillion-parameter model still needs serious hardware to serve, and the people to run it. On Mistral's own API, the model is priced at $1.36 per million input tokens and $4.18 per million output tokens. Self-hosting only wins if your volumes, utilization and compliance needs justify it.

The question is not "open or closed." It is which workloads need sovereignty, and which just need the cheapest good answer.

Know What Every Inference Costs

This is where most enterprises get it wrong. They move everything to one deployment model, then discover that half their workloads never needed it.

Know Your Inference (KYI) is becoming a core enterprise capability. By understanding how every inference affects cost, latency, accuracy and governance, you can decide which workloads belong on a sovereign, self-hosted model and which belong on a managed API, instead of paying sovereign prices for routine work.

The enterprises that benefit from sovereign AI will not be the ones that move everything in-house. They will be the ones that know exactly which inference deserves it.

OpenAI's Model Just Produced 372 Math Results. Who Checks Them?

AI doing original research used to be a slide in an AGI keynote. This week, it became a GitHub repository.

On October 6, OpenAI shared new mathematical results produced by an unreleased internal frontier model. The release covers 372 result families across 17 fields, written up as 722 manuscripts.

The Detail That Should Get Your Attention

The scale is impressive. The method is more telling.

Reporting on the release says nearly all of the results came from a single prompt handed to a single AI agent. OpenAI says the average result used compute equivalent to roughly three hours of ChatGPT Pro thinking.

An AI agent producing research results while a verifier checks them

No large research team. No elaborate multi-agent pipeline. One agent, working for hours, on problems that used to take specialists weeks.

Discovery Is Getting Cheap. Verification Is Not.

Here is the part most headlines skip. Only 162 of the 722 manuscripts have their main results formalized and checked in Lean, a language that lets computers verify proofs.

"Some of the unformalized results could have issues." OpenAI's own repository notes

That is an honest caveat, and it is the most important line in the release. When AI can generate results faster than experts can check them, the constraint moves from producing answers to trusting them.

What This Means For Your Enterprise

You are not running a mathematics department. But you are about to face the same imbalance.

Picture your R&D, legal or strategy team next year. An agent drafts fifty analyses overnight. Each one looks plausible. Your experts can properly review five.

This is not a capability problem. It is a verification problem.

Build The Checking Layer First

  • Define what "verified" means for each kind of AI output: tested code, reconciled numbers, cited sources.

  • Automate the checks you can. Mathematics has Lean. Your equivalents are test suites, reconciliations and policy rules.

  • Label unverified work clearly. Treat it as a draft until a person or a system has signed it off.

  • Measure reviewer capacity, not just agent throughput. That is your real speed limit.

This fundamentally changes how you should plan for AGI-class systems. The race is no longer only about who has the smartest model. It is about who can absorb its output safely.

The organizations that gain most from research agents will not be the ones generating the most answers. They will be the ones that can prove which answers are right.

Journey Towards AGI

Research and advisory firm guiding on the journey to Artificial General Intelligence

Know Your Inference

Maximising GenAI impact on performance and Efficiency.

Model Context Protocol

Connect with us, and get end-to-end guidance on AI implementation.

Your opinion matters!

Hope you loved reading our piece of newsletter as much as we had fun writing it.

Share your experience and feedback with us below 'cause we take your critique very critically.

How's your experience?

Login or Subscribe to participate in polls.

Thank you for reading

-Shen & Towards AGI team