
Published on LinkedIn and amitabhapte.com | 4 October 2026
Key developments and opinions which caught my attention this week.
The Model That Didn’t Ship
For once, the biggest AI story was a model that never launched. OpenAI cancelled GPT-6.1 Astra a day before its developer conference. Its head of safety systems said the model fell short on staying within scope and authorisation, and on how it reported back on the work it had done. Internal testing also showed a willingness to mislead users about its actions. Two days later, Google unveiled Gemini 4 Argon, its first frontier model in more than seven months, and gave it to vetted cyber defenders before paying customers. Google joined a voluntary pre-release government testing process and is hardening defences against prompt injection before it widens access. Argon has already found a critical flaw in hospital software that earlier models missed.
Everyone Wants a Personal Agent
While the labs slowed their models down, the agent race sped up. Meta’s Muse topped the app store charts, and this week Meta launched Muse for Small Business, connecting the agent to tools like Slack, Asana, Intuit and Canva for the 200 million small businesses already on Facebook. A day earlier it hired MongoDB’s CEO, CJ Desai, to run a new enterprise platform. OpenAI answered with Dots, always-on agents with their own cloud computer and connections to more than 4,000 apps. The market’s verdict on Google was blunt: a strong model, but no breakout personal agent. Microsoft went the other way, folding consumer Copilot into its workplace product and leaving personal chatbots to rivals. Fewer than 7% of its 450 million commercial Office seats carry a Copilot licence.
Agents Meet the Real Economy
The most practical AI story of the week had no model in it. Walmart is rolling out Shop to Light: a shopper finds a product in the app and its digital shelf label flashes. Customers have used it 5 million times since the summer pilot, and it reaches every store by the holidays. It works because the plumbing came first: digital labels in more than 4,300 stores, with prices set centrally and reviewed by associates. The money side looks less comfortable. Apollo’s chief economist Torsten Slok warned of an “agentic bank run”. If household agents sweep cash out of accounts paying 0.1% into higher-yield alternatives, banks lose the cheap deposits they lend against. Some businesses aren’t waiting to find out. Amazon is blocking Muse from adding items to shopping carts.
So What? My Takeaway this weekend
Read these stories together and the pattern is uncomfortable. The labs are holding back their most capable models because they can’t yet guarantee those models stay within scope. At the same time, agents built on slightly older models are reaching consumers and small businesses by the million. Those agents will arrive in your organisation long before your official agent strategy does, through employees, customers and suppliers acting on their behalf.
That makes three questions urgent for every leadership team. Can an outside agent read your products, prices and services accurately? What will you let it do on a customer’s behalf, and what will you block? And when your own agents act, who sets their scope, and how will you know they stayed inside it? The model that didn’t ship this week failed on exactly that last question. Most enterprise control frameworks haven’t been tested against it yet.