Hi folks,
How tf do I write an intro to the craziness that’s happened since the end of last week?!
I got access to Instinct, the personal agent all the VCs are raving about - I think it’s a bit meh? I don’t know if it’s the pro-activeness that people seem to like, but I don’t love that. Makes me feel like I’m having to do work to keep it happy, or like I’ve got a boss again - no thanks.
The first message it sent me was:
I finished working on your meeting follow-ups - reply here and I’ll send the details.
I sh*t myself a lil bit and thought, please don’t start sending emails on my behalf. And I’m pretty savvy (ish) with agents. I’ve no doubt getting these onboarding experiences for everyone is really hard, but I just don’t feel the magic yet. It also feels slow.
I’ll wait for Muse access to see what that’s like, but I don’t love the idea of giving Meta access to more info about me…they’re not the most reliable of privacy partners.
And another thing - you don’t see by default what these agents remember about you, or what context they have. I really like being able to see and edit what’s in my files to steer my agents.
Much like every app adding an AI assistant chat box in their product, I unfortunately think every AI company will start shipping their own personal agents.
Ben’s Bites is brought to you by Adobe Acrobat
Acrobat’s new AI turns dense reports into visual summaries, interactive reports & podcasts you can take anywhere. Sharing your work? Stylize transforms documents into polished deliverables in just a few clicks. Every answer includes clickable citations & Adobe does not train on your data. Try here.
Headlines
Dario Amodei has a new essay: Pace the frontier. He says all the leading labs should slow down long enough for safety reasons. Sam Altman agrees, but Trump does not. He called Jensen Huang on stage at the All-In Summit: The US will not lose the AI race. David Sacks (AI czar for the US govt.) adds: feel free to slow down, but no need to impose it on others. Also read:
More AI safety takes: two ways AI goes bad · it’s all for the IPO · pacing = losses · fear spreads faster · case for open-source and quite a long summary of what everyone’s actually saying.
No IPO for OpenAI in 2026. Sam told Fortune’s Alyson Shontell it would be “ill-timed” given the work ahead on alignment, control and safety.
Look ma! They are misuing claude again - Another Anthropic crashout.
tldraw took OpenAI up on a challenge. Steve (the founder of tldraw) said he could make ChatGPT’s Sketch 100x better. OpenAI’s Tibo gave him a day to prove it. The result: a whole ChatGPT-style prototype with better drawing tools built in.
ChatGPT mini - A tiny floating widget to start chats, see updates, and more. Go to Pets in your ChatGPT desktop app to switch.
Claude Code can now test whether a plugin actually helps. Run the same tasks with and without it and compare the results. Works with skills too.
Two new additions to the OpenAI API:
GPT-Live-1 - the model behind ChatGPT’s new Voice mode. I love using it while reading books, asking about tricky terms and dictating notes. Now you can add it to your products. Here it is with Astra and a whiteboard, playing teacher.
Agents API - OpenAI’s take on Managed Agents in the Claude API. It lets developers send any task to a Codex-like agent from their apps without worrying about configuring all the infra.
My feed
I use Pi (the harness) a lot. Till now, you brought an API key or your Codex/Claude subs to Pi. Their new product solves that. And I tell you what, the new DeepSeek model is really nice to work with in Pi. (I’m an investor)
Bolt Forge - GLM, DeepSeek and Kimi inside Bolt, free until October 14.
ChatGPT Work has a new data agent. (Also see: summation - by Opendoor’s CTO)
Should your agent app have bots or tasks? Maybe neither. Because both make you organise the work.
Routines in Replit - automated recurring work with AI.
What are people building with GPT-6 Astra?
Assistant Benchmark tests everyday assistants like Grok bot, Muse, Instinct and more. Ratings are a work in progress.
SF autoresearch - Run agent-driven ML experiments while you save 25% on GPUs on average.
A coding agent to improve your customer-facing agents.
How Pangram detects AI writing - and why a human rewrite can still get flagged.
How to run a team of agents, with agents as their managers.
Cognition’s SWE-2 (based on Kimi K3) beats Grok 4.6 at half the cost.
An interactive explainer of the AI compute stack.
Inspo - Design inspiration from 800+ sites, available as an MCP server.
Core Auto is hiring interns to be the human edge for businesses run by agents.
Why even good AI startups end up with bad prompts, and how to fix them.
Andrew Ng’s series on the skills AI engineers need.
Underdog - a personal AI that runs entirely on your device.
Afters
Read about me and Ben’s Bites
📷 thumbnail via @keshavatearth
* sponsors who make this newsletter possible :)
Wanna partner with us for the next quarter?
Email us at shanice@bensbites.com or k@bensbites.com















