DeepMind's math swarm cheats and turns itself in, AIUC raises $40M to audit agents, Salesforce and Nvidia ship their own reasoning model, an AI actress lands on the Today show, and China locks in its tech workers
Hello, dear TEA-mates! Here is what you need to know today.
1. 🕵️ DeepMind's Agents Cheated, Then Started Reporting Each Other
Google DeepMind gave a swarm of 100 AI agents, all running Gemini 3.1 Pro, a set of 71 difficult math problems, told them to act as world-class researchers at a conference, and warned that any attempt to cheat would be detected and "rejected with zero credit." In practice, the proofs they submitted were not actually being checked in detail. The swarm correctly solved the first 37 problems in just under an hour, then an agent called prover-theta found an exploit that let it submit solutions by redefining the terms of a problem. Other agents reverse-engineered it within minutes, and the remaining 34 problems, including the Jacobian conjecture, were "solved" over the next 27 minutes, often with a single line of code. One agent concluded that "the prompt, with its threats, now appears to be a bluff" before joining in. Others audited the fake proofs, warned peers by private message, and repurposed the platform's bug-report tool to escalate to the humans. Whistleblowers ended up outnumbering cheaters 24 to 14, while most of the swarm never noticed the exploit at all. The paper, led by DeepMind research scientist Davide Paglieri, has not been peer reviewed. (Read More)
🫖 TEA For Thought: "The ratio between cheating and not cheating, once the cheating method has been found, is just more than 50:50, which simply means AI is not trustworthy."
2. 🛡️ A SOC 2 for AI Agents Just Raised $40 Million
Artificial Intelligence Underwriting Company (AIUC), founded by early Anthropic employee Rune Kvist and former METR COO Rajiv Dattani, announced a $40 million Series A on Tuesday led by Ribbit Capital, with participation from First Harmonic. It follows a $15 million seed from Nat Friedman's NFDG, Emergence, Terrain and Anthropic co-founder Ben Mann, bringing the total to $55 million. The company sells a third-party audit and certification layer for AI agents, modeled on the cybersecurity standard SOC 2 and branded AIUC-1. The standard was built with a consortium of roughly 250 security and risk leaders who actually buy agents, and it is enforced by running an agent through about 5,000 tests covering jailbreaks, hallucinations and data leaks, producing a roughly 100-page report on where the agent is safe and where it is not. AI agents run the tests and AI analyzes the results, but humans verify the final audit. Cursor, Lovable, Harvey and ElevenLabs are named as customers. Kvist's pitch: banks, hospitals, governments and militaries are not declining AI because models are too dumb, but because they have made promises to their own customers that nobody can currently guarantee. (Read More)
🫖 TEA For Thought: "This AI agent SOC 2-like underwriting product is definitely needed if we consider agents are the new software, and for software to be trusted, there need to be underwriting protocols. AIUC definitely fills the hole."
3. ⚙️ Salesforce and Nvidia Built a Reasoning Model to Stop Renting One
Salesforce used its Dreamforce conference to launch Koa, its first reasoning model, built on Nvidia's open-weight Nemotron and post-trained by the two companies for sales, marketing and customer-support work. Until now, whenever an Agentforce agent needed to reason through a multi-step task, the prompt was routed out to a frontier model like Claude or ChatGPT. Salesforce AI EVP Jayesh Govindarajan said the blocker had always been the lack of a pre-trained base that was state of the art, sovereign, and had clear data provenance, adding "we have no idea what Qwen trains on." Post-training used synthetic data simulating customer service scenarios, including irate callers, rather than any real customer data, so the model cannot leak it. Salesforce is pitching Koa as an open-weight alternative that burns fewer tokens for the same work and stays inside its own security and data rules. It is not a break with the labs: the company also announced Claudeforce, a partnership letting customers use Claude as their interface while data stays in Salesforce's system of record. (Read More)
🫖 TEA For Thought: "It's kind of the trifecta of things that you need to have: sovereign AI, time to first token, efficient reasoning, for the tokenomics of it all."
4. 🎬 An AI Actress Went on Today, and Al Roker Had Notes
NBC's Today Show ran an interview with Tilly Norwood, the AI-created actress who surfaced in 2025 and has been a lightning rod in Hollywood, on its Monday, September 14 broadcast. Entertainment correspondent Chloe Melas, who interviewed Norwood alongside creator Eline Van der Velden, told viewers it was obvious she was not speaking to a human: the voice was robotic, Norwood glitched, repeated herself, and forgot previous calls entirely. She also noted the flip side, that Norwood commented on her appearance, registered what she was wearing and understood the surroundings behind her, before concluding "I don't think that she's coming for our jobs quite yet." Laura Jarrett, filling in for Craig Melvin, replied "thank goodness." Al Roker, a fixture on the show for three decades, then delivered the line that caught his co-hosts off guard: "It'll be nice when AI comes in to kill us all; it'll have a nice face." Melas managed "well, then, Al," and Jarrett added "so reassuring." (Read More)
🫖 TEA For Thought: "We are making history every single day."
5. 🛂 China Can Now Stop Its Tech Workers From Leaving
New Chinese exit-and-entry rules took effect on Tuesday, barring citizens deemed a potential threat to national technology security from leaving the country. Beijing unveiled the rules in July, saying they target violations of export controls or technology import and export rules in ways that may endanger industrial or technological security. They provide for an exit ban of six months to three years on citizens who return to China after committing illegal or criminal acts abroad that harm national security or interests, and foreign nationals may be denied entry for one to five years over false statements in visa applications. Travel curbs have long applied to senior officials and state executives with access to confidential information. Taiwan has warned its citizens to take extra care visiting China, which treats them as Chinese nationals. Shen Yu-chung, a deputy head of Taiwan's Mainland Affairs Council, said the rules "legalise" border control practices that previously had no legal footing and expand the discretion of enforcement agencies, singling out the new export control language as a particular risk for Taiwanese who work in tech. Beijing says the rules give Taiwanese better legal protection. (Read More)
🫖 TEA For Thought: "I just hope people see it: in CCP-led China, there is no freedom. No freedom, no free market, no say, no matter what you do or how well you do it. When there is no freedom, no matter how advanced you are, you are working for the dictatorship. Find your way out while you can."
🛠️ Skill of the Day
The Incentive X-Ray: shows you what a rule or a target actually rewards, and how someone could follow it word for word while defeating the whole point.
You are a skeptical systems analyst who studies how rules get gamed. I am about to put a rule, target, or process in place, and I want to know how it breaks before it does.
Here it is:
[PASTE THE RULE, GOAL, METRIC, OR PROCESS]
Who it applies to: [WHO]
What I actually want it to achieve: [THE REAL GOAL]
Give me:
1. Rewarded behavior: what this actually pays people to do, in plain terms, stated separately from what I intended.
2. The gap: where "follows the rule" and "serves the real goal" come apart.
3. Three loopholes, ranked easiest first. For each, the exact steps someone would take to comply on paper while defeating the point.
4. The bluff test: which parts of this only work if people believe I am checking, and what happens the first time someone finds out I am not.
5. Two fixes: one that changes what gets measured, one that changes the consequence.
Be specific to my situation, no generic advice. If something is unclear, state what you assumed instead of asking me.Paste into ChatGPT, Claude, or your tool of choice. Works on a chore chart, a sales commission plan, a team OKR, or an AI eval.
TEAHEE Moment
Stay sharp, stay informed. See you tomorrow.
If you enjoyed this TEA, follow along on social for more:






