Good morning. Gamers finally have a career plan. The skills that help them manage chaos on screen have caught the Transportation Department’s attention. It wants them to become air traffic controllers. The FAA was roughly 3,500 controllers short, leaving exhausted staff working overtime and flights delayed. So it ran a Fortnite ad telling gamers, “You’ve been training for this.” More than 2,000 signed up, prompting Transportation Secretary Sean Duffy to call it a “GAME CHANGER.” America’s aviation crisis now depends on whether they survive the tutorial.
In today’s dose:
- AI robots image control
- Agents fail data science
- Washington’s open model mystery box
- Self improving agent loop
- Grok 4.6, ranked
Researchers found a way to control AI robots with a picture. They hid instructions inside an ordinary looking image, so people see a picture while the robot sees a command. To test it, they placed one beside a robot trying to put a slice of bread into a basket. One image made the robot freeze. Another pushed its arm forward instead. Remove the image and the bot went back to work. The image only needed to cover 2% of the robot’s view to control it 77% of the time. At 5%, the attack worked 99% of the time. Security teams now have to worry about everything a robot can see. When the robot uprising begins, maybe we can stop it with a Picasso instead of an EMP. (arXiv)
The best AI agent fails almost half of real data science jobs. Researchers tested 15 agents on 275 jobs that asked them to turn raw data into answers a business could use. A typical job might ask the agent to study sales data and build a forecast. Claude 4.6 did best, finishing 57% of the jobs. GPT 5 reached 30%. Every open model came in below 1%. Many agents started the work and lost the thread before reaching the final answer. Anyone who has asked AI to analyze company data has probably watched it invent the answer. By comparison, two applied scientists and a master’s graduate completed 85% of the same jobs. Agents can speed up analysis, but check their work before making a decision. (arXiv)
Stop hallucinations. Power your agents & chatbots with real-time data from the Brave Search API:
- Specialized endpoints built for LLMs
- 40B page index, no Google scraping
- $5 per 1,000 calls, plus free monthly credits
- RAG pipelines, Claude MCP, OpenClaw, Hermes, and more
👀 closer look
Open models may soon need Washington’s blessing. The White House has figured out how to regulate AI without calling it regulation. Its new safety framework asks US labs to submit their best closed models for government testing before release. Nobody outside the government can see the framework, and the White House has no plans to publish it. Now officials want open models pulled into the same mystery box. Open source labs could volunteer to delay a model launch by 30 days while it undergoes testing by leading AI labs. But why would they want to? A government seal could make companies think the model has Washington’s approval, even though the framework carries no legal weight. Or they can skip Washington’s secret test and launch on time. Good luck convincing China to wait 30 days. (WIRED)
AI agents are learning how to improve themselves. LinkedIn built a customer support agent that learns from its mistakes. After the agent answers a question, other AI models check its work and show it where it went wrong. The agent uses that feedback to rewrite its own instructions, then tests the new version on past customer conversations. During a two week test, this loop helped the agent solve 27% more questions without updating the GPT 4o mini model underneath it. We’re nowhere near continuous learning yet, but shrinking weeks of manual work into hours or days brings it closer. For once, your angry support ticket may actually improve customer service. (arXiv)
Cursor may have rescued Grok from irrelevance. Grok had fallen behind OpenAI and Anthropic. Then SpaceXAI bought Cursor and gained access to its huge collection of real coding sessions. That data helped train Grok 4.6 to pause during longer jobs, check its work, and get back on track when something goes wrong. Its biggest gains came on those longer tasks, pulling Grok back near the leading models. Independent testing ranked it third overall alongside GPT 5.6 Sol. Now xAI needs people to use it. Just 4% of companies paying for AI tools currently use xAI. Cursor taught Grok how to recover from mistakes. Musk could use the same training. (Gizmodo, VentureBeat)
fun stats
🍽️ 87.5%. US venture funding that went to $100 million+ megadeals in the first half of the year. AI drove most of it. Everyone else is fighting over scraps.
✨ 1 billion. Active monthly users of Google Gemini app, making it Google’s fastest growing product ever.
🪼 94%. How much microplastics swarms of tiny cleanup robots can remove from water in just one hour.