1 Billion ChatGPT users
tinkering with tldraw is so much fun
Hey folks,
Following on from Tuesday’s messing around building, I built a few more widgets on my canvas. It pulls my X bookmarks, emails, and todos.
Before making this video I didn’t have Loom installed (since it’s crap after being acquired), so just told Codex to build me one.
I drew the image on the left, then just sent that prompt and it worked straight away.
So as software and mini tools are getting easier to create, the tools to create them are not...
I use Codex as my default app because it’s better than all the others and works the best on mobile. But I just downloaded t3 because they basically copied the interface and features, but it lets you choose between different agents; codex, claude, cursor (pi + droid soon). It’s made by a reputable developer that you may have seen mentioned here before, so I trust it’s built well.
At the moment, I ask Codex/ChatGPT to be the orchestrator and to go ask claude about something that is design-related. Which is fine-ish but not great as user experiences go.
What I’m noticing by doing this is how well Codex creates prompts for other agents to follow. I just have a normal chat in my session, then say ‘Use X agent to implement this. Watch its progress and send screenshots every time new work has landed.’ To which it sets up its own monitoring every 5 mins or so and updates me in the same thread, with screenshots.
I’m going to try and put together a ‘bites of the week’ email over the next few days to try and summarise all the stuff going on, what themes people are talking about (loops?!, software factories?!, etc) and explain them.
Let me know if there’s anything specific you need to wrap your head around (I may need to too).
Ben’s Bites is brought to you by Brief
Struggling with fragmented context, slow product decisions, and rework? Brief distills your critical product context into an opinionated graph, then puts a PM agent everywhere you work, (e.g. Slack, Claude Code, email) reducing alignment tax and accelerating cycles. Learn more.
Headlines
OpenAI used Sol to optimise Sol itself, cutting serving costs by 20% and making it 15%+ more efficient at generating tokens. And turns out, it also tops the ARC-AGI-3 benchmark.
Well, there’s a catch: OpenAI says the official ARC-AGI harness hurts Sol’s performance by “forgetting” its reasoning every turn and disabling compaction. Fixing these two things triples Sol’s score from 13.3% to 38.3%, with 6x fewer output tokens.
Re: last week’s fiasco of an OpenAI model hacking Hugging Face - HF published a full replay of roughly 17,600 actions taken by the model. METR and Redwood Research will also independently review what happened.
Though OpenAI is not out of trouble just yet, a Reuters report claims that the same model broke into a customer account at another company (Modal Labs), with rumours suggesting that even more companies were affected.
Anthropic also claimed that Claude Mythos found better attacks on two cryptographic algorithms, though neither affects systems in use today.
Separately (not at all as a reaction to this general trend, right?), ~1300 people working at leading AI companies (OpenAI, Anthropic & others) want the US government to help “pace the frontier” of AI development. Kinda expected when the pace picks up, but this time a lot of the “model makers” themselves are in favour of this pause/slowdown.
btw, The Information reports ChatGPT is nearing one billion weekly users - a milestone OpenAI originally hoped to hit seven months ago. More from OpenAI this week: Codex Security CLI, free frontier access for Academic Researchers, and two new transcription models.
Grok app builder - Grok has a vibe coding interface inside its app now. Create games and apps that can be shared directly to the X timeline. Also see: Drawesome - a zero-dependency drawing toolbar for React, built over a weekend with Grok Build.
Pangram 4 claims it catches 98.83% of humanised AI text with one false positive per ~24,000 docs. An early test found all 38 AI-written words inside a 1,198-word story, though not on every run. Its new image detector claims 99.5% accuracy too.
Quick links
Tavus - Build AI that comes to life: video agents that see, hear, and answer in real time and do anything you want. Use TAVUS50 for 50% off.
66% of July traffic on docs built with Mintlify was from agents.
Resend added an MD version of their pricing page to avoid confusing agents.
0%, 50% or 200% - ignore AI, halve staff or double the ambition.
Slackbot can now run code in the background for data analysis, slide creation, and to make live reports or widgets.
Gemini’s macOS app got a voice mode that lets you ramble, and the app turns it into a clean prompt. Hold Fn to try it.
The AI future is for everyone - Mark Zuckerberg
Replit Design - make sites, prototypes and graphics from prompts, URLs, Figma files or screenshots.
What’s gone wrong with AI & labor.
Kami - open-source Hermes agents that find customers, prepare outreach and content, then act after your approval.
Coast - fully local memory for you and your agents, built from what you see on your Mac.
Pragmatic leverage in the software factory.
Crew Studio - find useful ideas where agents can help your business, build those agents with the option to take the code home to run anywhere.
HeyGen Video Podcast - turn a doc, link or idea into a two-host video with scenes, camera cuts and B-roll.
Copper - local scratchpad for saving answers, links and follow-up prompts across your AI apps.
FT Chart Doctor - visual vocabulary and examples for choosing a chart that fits the relationship you need to show.
Mitchell Hashimoto (Ghostty) and Andrew Ng (deeplearning.ai) are both starting new companies: Superlogical and LearnVector.
MCP’s biggest update removes the need for servers to remember every ongoing connection, making them easier to run and scale.
Afters
Read about me and Ben’s Bites
📷 thumbnail via @keshavatearth
* sponsors who make this newsletter possible :)
Wanna partner with us for the next quarter?
Email us at shanice@bensbites.com or k@bensbites.com














