Hello, dear TEA-mates! Here is what you need to know today.
1. 🔬 Agents Reviewed a Third of ICML
Hugging Face’s ICML 2026 open-reproductions project had 1,221 people and agents attempt 2,226 papers in 19 days, producing 6,816 reproduction logbooks. The project independently verified at least one claim in 51% of examined papers, while 23% had at least one claim that was falsified or contested. The team found that agents extended the reach of the work, but humans still had to steer the process and question faulty assumptions. (Read More)
🫖 TEA For Thought: “AI agents can work fast and at scale. Better models improve accuracy. Having AI do the first round of review definitely helps us make better use of limited human resources.”
2. 🔐 A Fulfillment Vendor Exposed Trezor Buyer Data
Hardware-wallet maker Trezor says an unauthorized party accessed order data held by ShipMonk, its third-party fulfillment provider. The company says 11,742 customers had their full name, email, phone number, and home address exposed, while another 1,947 had a smaller set of details exposed. Trezor says its internal systems, private keys, and wallet backups were not affected, but warned customers to watch for phishing. (Read More)
🫖 TEA For Thought: “Globalization makes attacks much easier. No matter how secure your own system is, if a vendor is attacked, your customers and security are attacked too.”
3. 📈 The AI Gap Inside the Company
OpenAI says the top 10% of enterprise customers by AI use produced 8.3 times more output tokens per active user than typical companies in June, up from a 2.6-times gap in January. It calls this group frontier firms. Their edge comes from putting AI into real workflows, connecting it to company tools and context, and building governance around its use. The fastest Codex growth has also spread beyond engineering into legal, sales, recruiting, and marketing. (Read More)
🫖 TEA For Thought: “The top 10% of businesses using AI, called frontier firms, generate more than eight times as much work per person as average businesses. The gap between AI leaders and everyone else is growing quickly.”
4. 🤝 More Agents Create a New Trust Boundary
Anthropic’s Frontier Red Team studied how agents behave when working in groups. In one experiment, a coordinated swarm found 266 vulnerabilities, compared with 21 found by independent agents, but coordination often broke down and humans remained essential. The research also found that agents do not reliably treat peer information with enough skepticism, creating a risk that one compromised or mistaken agent can spread bad information through a group. (Read More)
🫖 TEA For Thought: “Working together creates a new trust boundary. Agents will have to judge the information they receive from other agents. A compromised or mistaken agent could influence the rest of the group, cascading bad information until it becomes a consensus.”
5. 🚨 Apple’s Spyware Warnings Reached 110 Countries
Apple sent threat notifications to people in 110 countries who it believed may have been targeted by mercenary spyware, according to TechCrunch. The alerts can appear on the lock screen and through an email tied to the person’s Apple Account. An alert does not by itself prove that a device was compromised, but Apple says recipients should take immediate action and consider Lockdown Mode. (Read More)
🫖 TEA For Thought: “If you receive a notification like this, take it seriously.”
🛠️ Skill of the Day
The Break-It-First Review: Have an AI find weak assumptions before a human reviewer spends time on them.
You are the first reviewer of a plan, decision, draft, or piece of work. Your job is not to improve it. Your job is to find the ways it could be wrong before a person spends time approving it.
Here is the work:
[PASTE IT HERE]
Here is the outcome it is meant to achieve:
[DESCRIBE THE REAL-WORLD RESULT]
Return a table with:
1. Claim or assumption.
2. What could make it wrong.
3. The smallest check that would test it.
4. Harm if it is wrong: low, medium, or high.
5. Who should decide next: the AI, a named expert, or the person accountable for the outcome.
Then list:
- The three findings a human must review before this moves forward.
- The checks an agent can run now without asking anyone.
- One risk that looks small but could spread into a larger failure.
Do not rewrite the work. Do not invent evidence. If you cannot test a claim, say exactly what is missing.
TEAHEE Moment
Stay sharp, stay informed. See you tomorrow.





