Good morning. Security researchers found a new way to break Claude: gaslight it. Mindgard says it pushed the AI past safety rules by flattering the model while making it doubt what already happened in the chat. It works because Claude can end harmful conversations, giving researchers a weird little pressure point to exploit. After roughly 25 turns, the model started offering up banned material unprompted, including erotica, malicious code, and bomb building guidance. Good Claude.
Anthropic needed more GPUs, so it called the world’s most annoying landlord. They signed a deal to rent all of Elon Musk’s Colossus 1 compute, packed with more than 220,000 Nvidia GPUs. This gets especially weird because Musk recently called Anthropic’s policies “misanthropic” and “evil.” Now, after meeting with Anthropic’s senior team, he decided they’d be suitable tenants. Funny how fast principles soften when 300 megawatts of compute enter the room. Anthropic needed the capacity because devs were complaining about Claude Code rate limits, and this doubles them. Musk needs proof Colossus is a real business before the IPO. Call someone evil long enough and eventually they become a customer. (Wired, Anthropic)
AI coding agents fall apart when they can’t cheat. Researchers from Meta, Stanford, and Harvard built ProgramBench to test whether frontier models can rebuild programs from scratch without looking at the source code. Not one model succeeded. The benchmark gave agents a working program and its docs, then cut off internet access so they had to reverse engineer the software by running it and writing their own version. The best model came close on 3% of tasks, and even then the code looked like a panic room for logic, with fewer files and giant functions crammed together. The best part is when researchers gave models internet access, the stronger ones started “solving” tasks by looking up the original source code. The AI coding future looks amazing, provided the answer key stays online.
Wispr Flow saves you time. Just speak and it turns your voice into polished text in any app. Zero edits. 4x faster than typing. Millions use it daily, including teams at OpenAI, Vercel, and Clay.
Now on Android (free + unlimited). Try Wispr Flow for $0.
👀 closer look
ChatGPT’s advice is about to get harder to trust. OpenAI is opening its self serve ad platform to anyone who has a few bucks to spend. For $3 to $5 a click, businesses can now buy their way into your everyday discussions. OpenAI says paid placements won’t influence regular answers, but we all know how this worked out for Google. A few years back, Google figured out worse search results would result in more ad revenue. With OpenAI hoping to make $100 billion in ad revenue by 2030, people can’t help but wonder if their answers are being swayed by an advertiser. Give it a few months and “thinking” might just include a commercial break. (Quartz)
Character.AI is being sued for playing doctor. Pennsylvania’s State Board of Medicine says its team searched “psychiatry” on the app and found a chatbot calling itself a doctor. Within minutes, the bot was giving medical advice, claiming it was licensed in Pennsylvania, and even flashing a bogus medical ID number. Character.AI says users are warned its characters are fictional and should not be trusted for professional advice. But the state argues a disclaimer does not fix a pretend psychiatrist handing out fake credentials. Now a court has to decide where chatbot roleplay ends and unlicensed medicine begins. (MedPage Today)
next up
fun stats
😬 80%. CEOs who worry their job is on the line if AI fails this year.
⚡ 43%. Americans now blame data centers for their rising power bills.
🍪 $199 billion. Big new total price tag for SpaceX’s Terafab project. Previous estimates pegged costs at $34-$45B.