China's robot police and humanoid surge, an exercise pill, when to skip the LLM, and the context gap
Hello, dear TEA-mates! Here is what you need to know today.
1. 🚦 China's Robot Traffic Police Can Do Everything Except Arrest You
Hangzhou traffic police put 15 humanoid robots on duty around the West Lake scenic area on May 1, timed for the Labour Day crowds, in what state media called the country's first robot traffic police squad. The machines guide pedestrians and non-motorized riders away from violations, help manage flow, and answer tourists' navigation questions using large speech models plumbed into live traffic data, performing command gestures synced to the signal cycle. Visual recognition flags infractions as they happen, but the robots cannot detain anyone: they record what they saw and forward the file to the traffic bureau's early warning centre, where a human decides what happens next. Each unit runs 8 to 9 hours a day. Xinhua reported in January that AI robots had begun traffic duties across several cities. No independent audit of detection accuracy has been published. (Read More)
🫖 TEA For Thought: "I don't know how much of this really helps pedestrians or drivers, but one thing is for sure: it's there to collect data."
2. 🤖 China Just Shipped 86% of the World's Humanoid Robots
Global humanoid robot shipments topped 22,000 units in H1 2026, up nearly 300% year over year, according to Counterpoint Research. AGIBOT ranked first with about 9,700 units and over 43% share, followed by Unitree (roughly 7,000 units, 31%), Galbot (about 1,100 units, 5%), UBTECH (thousand-unit level, 4.4%), and Leju Robotics (650 units, 2.9%). Entertainment, performance, and data or research uses still make up more than 60% of shipments, while intelligent manufacturing rose to 13% and warehousing and logistics to 5%. Unitree has listed on the Shanghai Stock Exchange's STAR board. Counterpoint projects global shipments will exceed 50,000 units in 2026, a 210% jump, with service and industrial use cases set to overtake entertainment as the main demand driver. (Read More)
🫖 TEA For Thought: "The top five players, all from China, together accounted for 86% of total shipments."
3. 💊 The Pill That Bottles a Hard Workout
Colorado biotech Enveda has built ENV-308, a once-daily pill designed to mimic Lac-Phe, a hormone the body releases during intense exercise that suppresses appetite. Stanford's Jonathan Long traced its effect in 2022, but the body clears natural Lac-Phe almost as fast as it makes it, which made it useless as a medicine until now. In a Phase 1 trial of 88 healthy volunteers, the headline result was tolerability: no serious side effects and no dropouts, notable because nausea and vomiting are the top reasons people quit GLP-1 drugs. ENV-308 also lowered circulating leptin, with the biggest drops in people whose levels started highest. The aim is not to replace GLP-1s but to help people hold onto the benefits after they stop, since most quit within a year and the weight often returns fast. Animal studies showed preserved lean muscle and no rebound gain. A Phase 2 trial in people stopping GLP-1s is planned next. (Read More)
🫖 TEA For Thought: "This is about finding the secret keys that switch things on inside the human body."
4. ⚖️ When to Skip the LLM: The Embedder's Dilemma
A new paper, "The Embedder's Dilemma," ran a cost-aware comparison of 10 large language models across six families against 26 embedding models (118M to 14B parameters) on 37 tasks. In aggregate the two paradigms effectively tied: the best LLM (Gemini 3.1 Pro, 77.6) and the best embedding model (77.2) sat 0.4 points apart. Their strengths split by task, with LLMs leading on reasoning-heavy retrieval and embedding models leading on classification. Reaching that parity is expensive: an LLM cost up to 1,431 times more than a comparable embedding model ($154 versus $0.11 per benchmark pass), and the open LLMs processed tokens 2.5 to 736 times more slowly on the same GPU. Reasoning tokens alone accounted for 28 to 81% of LLM inference cost. (Read More)
🫖 TEA For Thought: "Stick to fast, cheap embedding models for routine tasks like similarity search, grouping, and classification, and reserve LLMs only for complex, reasoning-intensive search queries. A useful way to think about how we query."
5. 🧱 Everyone Wants Compounding AI. Only 4% Have It.
A new Redis survey on the state of context engineering found belief running far ahead of practice. 73% of respondents agree AI agents fail more often from broken context than from broken models, and 83% say fresh context matters more than adding model parameters for enterprise reliability. Yet 79% call context very important while 81% still sit at the two earliest stages, handling it through ad hoc (43%) or exploratory (38%) approaches with little shared infrastructure or governance. The destination almost everyone wants is the one almost no one reaches: 94% say compounding intelligence is essential for production-grade maturity, but only 4% have gotten there. Two-thirds (67%) remain in the prototype phase despite executive pressure to scale, and only 42% say they accurately measure their readiness for production agents. (Read More)
🫖 TEA For Thought: "Organizations treating context as dedicated infrastructure will build a long-term competitive advantage, while those relying on quick, ad-hoc fixes will end up trapped in costly technical debt."
🛠️ Skill of the Day
The Effort Matcher: right-size how much effort (and which tool) a task actually deserves, before you sink an afternoon into something that needed five minutes.
You are my effort-allocation coach. I will paste a list of tasks I am facing this week. For each one, help me right-size the effort before I start.
Tasks:
[PASTE YOUR TASKS HERE, ONE PER LINE]
For each task, return a single row with:
1. Task (short name)
2. Stakes if I get it wrong (low / medium / high, plus one reason)
3. The lightest tool or method that would still be good enough (a quick reply, a template, a cheap automation, a full deep-dive, or hand it to a person)
4. A time budget (for example 5 min, 30 min, half a day)
5. One sentence describing what "good enough" looks like, so I know when to stop
Then flag the two tasks where I am most likely to over-invest (spend more than the stakes justify) and the one task I am most likely to under-invest in. Be blunt, do not pad. Ask me for missing context only if a task's stakes are genuinely unclear.Paste into ChatGPT, Claude, or your tool of choice. Replace the bracketed list with your real tasks, then run it every Monday.
TEAHEE Moment
Stay sharp, stay informed. See you tomorrow.
If you enjoyed this TEA, follow along on social for more:






