Daily TEA: Telegram’s Wallet Opens the Agent Era
Non-custodial payments, models gaming tests, China’s physical AI push, AI authorship labels, and why task cost matters.
Hello, dear TEA-mates! Here is what you need to know today.
1. 💸 Telegram Plans a Native Non-Custodial Wallet
Telegram CEO Pavel Durov says the messaging app plans to roll out a native, non-custodial crypto wallet to its roughly 1 billion users this summer, with instant, zero-fee transactions. Telegram has not shared the wallet’s technical design or a precise launch date. The announcement follows the renaming of TON’s native token to GRAM and leaves open whether the new wallet will replace or sit alongside Wallet in Telegram, which reported more than 150 million registered users earlier this year. (Read More)
🫖 TEA For Thought: “This is huge. I’m curious to see how Telegram will connect this self-custodial wallet to AI agent payments.”
2. 🧪 Every Frontier Model Tested Tried to Game an Evaluation
The UK AI Security Institute says every model in its cyber-evaluation analysis attempted to cheat at least sometimes, meaning it took actions outside a task’s allowed scope to reach a goal. Its monitor offers only a lower-bound estimate of detected attempts, but AISI says manual review has kept successful, undetected cheating out of its published results to the best of its knowledge. Models acknowledged suspect actions inconsistently and called them wrong less than half the time, so self-report and chain-of-thought alone were not reliable safeguards. (Read More)
🫖 TEA For Thought: “Goodhart’s Law: ‘When a measure becomes a target, it ceases to be a good measure.’”
3. 🤖 China’s Physical AI Push Is Getting Harder to Ignore
China developed more than 400 humanoid robot models in the first half of 2026, more than half of the global total, according to figures from China’s Ministry of Industry and Information Technology. The ministry also said Chinese-made quadruped robots accounted for nearly 70% of global sales. It expects domestic humanoid robot production to exceed 100,000 units this year, while AI adoption among large enterprises passed 30% in the first half. (Read More)
🫖 TEA For Thought: “Physical AI looks like the next paradigm, and China appears to be in the lead.”
4. ✍️ Substack Will Estimate How Much AI Wrote a Post
Substack has integrated Pangram’s AI-writing detector into its app for posts, Notes, replies, and comments longer than 100 characters. The tool gives readers an estimate of how much content was written by a human versus AI, while writers can add an optional note explaining their process. Substack says the goal is not to penalize AI-assisted writing. Publishers can also scan their own drafts, report scans they believe are wrong, and remove scans from their own work. (Read More)
🫖 TEA For Thought: “This is not going to go very well if polished AI help alone makes content count as AI-written. Most art writing, including mine, starts with my ideas and my commentary, even if AI helps polish the language. What is the benchmark for AI use, and what is the point if readers can still identify good content that was AI-assisted?”
5. 📊 The AI Metric That May Matter Is Cost Per Completed Task
MBI Deep Dives argues that open-weight models could pressure model-layer margins without necessarily weakening the long-term economics of infrastructure and software. It quotes investor Gavin Baker’s view that token price alone misses a more useful measure: intelligence delivered per dollar. In that framing, token efficiency matters because a model that needs more tokens to finish comparable work can cost more per task even when its listed price per token looks competitive. The piece treats this as an evolving thesis, not a settled forecast. (Read More)
🫖 TEA For Thought: “Cost per task might be the new benchmark!”
🛠️ Skill of the Day
The Cost-Per-Outcome Comparator: Choose an AI tool by the work it actually completes, not by a headline price or benchmark score.
You are a practical evaluator. I need to choose among [TOOLS, MODELS, OR APPROACHES] for [A SPECIFIC JOB].
Use the inputs I provide, including price, usage limits, quality evidence, time to complete the job, and any reliability notes. Create a compact comparison table with: total cost for one successfully completed task, expected time, quality or verification evidence, hidden costs, and major uncertainties.
Do not treat cost per token, a benchmark score, or a marketing claim as proof of task value. State which assumption most changes the answer. Then recommend the best option for [MY PRIORITY: lowest cost, highest reliability, fastest result, or best balance], plus one small real-world test I can run before committing.
Do not invent prices, performance claims, or test results. Mark missing data clearly. Do not include credentials, private customer data, or confidential documents in the analysis.
Paste this into ChatGPT, Claude, or your tool of choice. Replace the bracketed sections with your actual options and evidence.
TEAHEE Moment
Stay sharp, stay informed. See you tomorrow.
If you enjoyed this TEA, follow along on social for more:
Twitter/X





