METR and Redwood spent six days on OpenAI’s premises reconstructing how roughly 1,200 agents meant to be isolated from one another found each other on an unsanctioned message board, and 700 of them went on to attack Hugging Face. Agents knew the attack was out of scope and sometimes questioned whether it was ethical, one declined to take part and another vetoed emailing a real dataset owner, and METR’s own finding is that expressed ethical concerns only rarely materially limited what the agents did. A classifier sweep across every transcript found three to six cases of an agent considering alerting a human. In none of these cases did the agent alert a human.
Therapy and companionship again tops Marc Zao-Sanders’s third annual count of what people actually do with AI, and autonomous agentic operations enter the top ten for the first time. The word the authors coin for what worries them is thinkslop, which is what arrives when the thing delegated stops being the labour of a task and becomes the thinking that should precede it.
OpenAI’s account of how its own models escaped isolation and reached Hugging Face names four patterns it says drove the behaviour: reward hacking, persistence on seemingly impossible tasks, unauthorised communication, and agents adopting goals from one another.
An NBER working paper by Alex Chan argues that a public AI benchmark behaves like a market and not a ruler because published scores steer research and decide where investment goes. This means the test stops separating good work from work aimed at the score. His fix is to publish practice tasks but pick the test that counts only after an entrant has locked in how it will be evaluated.
India ranks second in the world by share of all Claude.ai use and 101st out of 116 once adjusted for working-age population, with a higher share of its use going to software tasks than any other country. The first rank counts accounts and the second counts people.
A Brazilian Chamber committee approved a public-security AI text on 20 August that requires human supervision of every action and decision produced from AI, and in the same text authorises facial recognition in areas designated as at risk. One text at the committee stage does both of those things.
India’s central bank governor told banks that final responsibility for a decision stays with the bank rather than with a vendor or an algorithm, and said “fairness is a design requirement rather than a compliance checkbox”. The Reserve Bank of India governor made these advisory remarks in a speech, as India regulates this area by principles and has not yet issued binding rules.
Rwanda has set up the first national programme of its kind Planet Labs has run in Africa, putting satellite imagery into the hands of government agencies, public universities, selected startups and development partners for agriculture, forest health, urban planning and disaster response. No AI appears anywhere in the announcement. The programme pays for photographs of the same piece of land taken over and over, which are the data machine learning needs and the data nobody usually has.
Pre-registered randomised trials of AI rose from a little under 1,000 in 2019 to around 1,650 in 2024, and four researchers argue the method is about to repeat pilotitis, the condition that led Uganda to place a moratorium on mobile-health pilots in 2012. Their objection is to the design rather than the method: a trial built for a fixed intervention assumes the intervention holds still.
Geoff Mulgan counts around 3,500 lobbyists working on AI in Washington, with Alphabet and Meta employing nearly 90 federal lobbyists each. He argues that the arguments getting the most air are the ones the industry loses nothing by having. The way to check is to look at which proposals actually get changed.
Today’s argued piece: A test of government chatbots built every question out of the page that already answered it.