the Microdose

Bot Blooper Loop

+ Claude managed agents and Dr. Google
Adam Wildheart
Recursive AI founder Azalia Mirhoseini
Recursive AI founder Azalia Mirhoseini

Recursive AI founder Azalia Mirhoseini/The Microdose

Subscribe to The Microdose
Subscribe -- post
Cheri Wildheart
Adam Wildheart

Good morning. Imagine assembling a dream team of AI geniuses, giving them billions of dollars, and promising the world superintelligence. Then, when it’s time for the big reveal, you introduce an AI whose greatest strength is suggesting shoes. Meta’s new model, Muse Spark, might be an upgrade from previous embarrassments, but it still trails GPT and Gemini. While Anthropic worries its latest AI could end civilization, Zuck delivered an AI that poses zero threat to humanity or its competitors.

Agents get better by rehearsing their own stupidity. Stanford researchers, including Recursive AI founder Azalia Mirhoseini, got tired of AI agents making the same dumb mistakes. So they built TRACE, a tool that pinpoints exactly where the bots go wrong. It finds recurring problems like fumbling simple tool commands or citing imaginary laws, then makes agents rehearse these mistakes in a sandbox until they improve. And it works. Agents trained by TRACE scored 14 points higher on customer service tasks and nailed 7 perfect scores on tool use benchmarks, all while updating only a tiny fraction of the AI’s brain. Now instead of being consistently bad, agents can practice tasks in private until they get it right. Bring on the blooper reel. [arXiv]

Three words that will make you smile: Claude Managed Agents. Anthropic is rolling out a new service designed to take the pain out of deploying agents. Businesses can now launch secure, cloud based agents just by describing them, while Anthropic handles all the messy stuff like sandboxing and credential management. Anthropic says it speeds up deployment by 10x. Pricing is simple: regular Anthropic token rates, plus $0.08 per active session hour. Advanced features like memory tooling and multi agent orchestration are still in limited preview, but Anthropic already covers enterprise essentials like governance and execution tracking. Wake me when agents can learn just by watching. [New Stack]

There’s a shortcut to ship scalable AI agents. Google Cloud’s Startup technical guide gives you pre-built frameworks to design, build, and deploy intelligent autonomous systems faster.

Inside the guide, you’ll discover:

  • Frameworks to accelerate autonomous agent design from day one
  • Best practices for prompt engineering workflows
  • Frictionless deployment strategies using Google Cloud infrastructure

Get the guide

DARPA wants AI agents to invent their own language. The Pentagon has thrown down a $2 million prize for anyone who can create a mathematical framework for AI-to-AI communication. Right now, agents swap raw data without context, making collaboration tough. DARPA thinks giving them clearer ways to talk to each other could speed up scientific discovery, and even allow them to rediscover breakthroughs like the periodic table from scratch. If it works, the next step is letting AI agents invent entirely new scientific ideas using their own math based language nobody else understands. The machines can finally talk science without having to dumb it down for the rest of us. [The Register]

Anthropic’s fight with the Pentagon just hit a major setback. A federal appeals court in DC ruled Anthropic must keep its supply chain risk label, firmly siding with the Pentagon. The decision directly contradicts an earlier court ruling in California that the Pentagon likely acted in bad faith. Anthropic argues it was targeted after setting ethical limits on how Claude can be used. But the appeals court said removing the label could disrupt critical military operations, leaving Anthropic stuck in legal limbo. Experts say Anthropic’s case is strong, but courts rarely overrule the Pentagon on national security. [Wired]

Reach your audience

Get in front of 70,000+ tech decision makers with budget and technical influence. 95% US based. 66% open rate. Read daily by AI leaders and builders at high growth companies. 

Does cyber insurance cover AI agent screw ups? Researchers from Microsoft, Google, and Columbia University just proposed a financial risk framework. They argue technical safeguards alone can’t stop AI agents from occasionally blowing things up – especially when real money’s involved. ARS sets up escrow accounts and underwriting to pay you back whenever an agent goes rogue. Simulations showed ARS dramatically cut financial losses, making it safer to trust AI with high stakes tasks. Handing your wallet to AI is easier when someone else covers the cleanup. [arXiv]

Google’s new AI nuked your healthcare startup’s moat. MedGemma 1.5, Google’s multimodal med AI, is fully open source and built for serious clinical tasks. It handles everything from 3D radiology and detailed pathology scans to analyzing chest X-rays and finding key anatomical details. MedGemma also extracts structured data from dense lab reports and electronic health records, translating them directly into industry standard formats. Recent benchmarks show it boosted accuracy in pathology by 47%, chest X-ray localization by 35%, and record analysis by 22%. On Q&A tasks, it hit 89.6% accuracy. Google just gave healthtech and biotech builders a fully open med AI stack. Time to update those investor slides. [Hugging Face, arXiv]

fun stats

🌐 $33 trillion. Money moved by stablecoins last year, according to Morph – more than Visa and Mastercard combined. 60% came from B2B transactions, as companies used dollar tokens to shift money globally. 

🔧 44%. GenZ workers who are actively trying to sabotage AI rollouts, fearing job loss. In a recent news poll, just 26% of registered US voters think AI is great. 

🧠 #1. China has surpassed the US in the number of top AI researchers. If the trend continues, by 2028 top Chinese elite AI talent could outnumber US-based ones by 2 to 1.

subscribe today

Get an edge with The Microdose

Skip the prompts, get the signal. We’ll send you short daily updates about the real AI + future tech you need to know. Fast, smarter, mildly addictive – and free. 

Subscribe -- page