Week 31 in AI: models can move, work memory cannot

Models are becoming easier to swap. The work memory around them is not.
That is the strongest signal from this week. Perplexity is gathering files, instructions, Skills, and connectors inside Projects. Google is making memory and agent identity part of its agent platform. Claude Code Routines lets work continue in the cloud under linked GitHub identities. OpenAI is making Codex sessions persistent and forkable.
If that direction holds, the next major AI lock-in will not be a model name. It will be everything the tool remembers about how you work.
The model can move while the context stays behind
Claude Opus 5 appeared this week through Anthropic, Google Cloud's Model Garden, and Perplexity. Perplexity also launched Gateway API, which puts models from several providers behind one key and familiar API schemas. MCP 2026-07-28 removed protocol sessions and made the server layer more stateless, cacheable, and routable.
That is a clear counterforce to model lock-in. Intelligence itself is becoming more interchangeable.
The work around the model is becoming more persistent. It includes files, decisions, reusable instructions, connected accounts, execution history, and knowledge about who may do what. Changing a model may soon be a configuration edit. Moving three months of working context is a different job.
Six vendors are building six forms of work memory
Anthropic Claude introduced Claude Code Routines as a research preview. Saved prompts can run on a schedule or start from API and GitHub events on Anthropic-managed infrastructure. The automation no longer depends on a local laptop staying awake, but the routine becomes more closely tied to the host platform and its connected identities.
Google Gemini made Agent Runtime, Memory Bank, and Agent Identity more widely available. A run can continue for up to seven days. Structured facts can persist across interactions, and an agent can have its own lifecycle-managed identity. Gemini Spark also started using signed-in Chrome sessions and saved credentials with the user's permission.
Perplexity turned Spaces into Projects. Shared files, instructions, Skills, automations, connectors, and channel bindings can persist as one workspace. It may be the clearest example this week of an AI product trying to become the place where work memory lives, not merely the place where a question gets answered.
OpenAI released Codex 0.146.0 with persistent and forkable sessions, Agent Plugin manifests, and better recovery after interrupted work. Its first stable Terraform provider also provides a way to describe projects, identities, roles, limits, and settings as code.
Mistral continued building versioned Prompts and Skills alongside durable workflows. Vibe 2.23.2 added guided skill creation and showed the origin of each configuration value. Once behavior has versions and provenance, it becomes an asset a team can improve—and something it needs to be able to move.
xAI/Grok combined resumable workflows with Grok Build, publishing from chat, and browser automation through TinyFish. Grok also moved into Google Workspace and GitHub Copilot. Its memory is less concentrated in one project container, but the direction is similar: the model is being attached to more surfaces where work and logins already live.
Lower model prices make work memory more valuable
OpenAI cut the API price of GPT-5.6 Luna by 80% and Terra by 20%. Sol gained Fast mode for output up to 2.5 times faster at twice the Standard price. In the other direction, xAI's new Think Fast 2.0 voice model raised list pricing from $0.05 to $0.08 per audio minute.
Model and execution prices can move quickly. A cheaper model can take over one workflow step in an afternoon. The accumulated value in project files, routines, connectors, and history does not move as easily.
The economic test should therefore cover more than token pricing:
- Time to an accepted result: not merely cost per request.
- Value of reused context: how much work does the team avoid repeating?
- Migration cost: can files, instructions, history, and permission logic be exported in useful formats?
- Interruption cost: what happens to active work when the central service is unavailable?
The strongest counterargument: open layers may win
This week's developments do not point only toward lock-in. MCP's new stateless architecture, Perplexity's Gateway API, and Claude Opus 5 inside Google Cloud show that models and tool servers can become easier to replace.
That could limit vendor control. An organization can keep persistent context in its own document, database, and identity systems, then use several models above them. The AI provider becomes an engine rather than the archive of how the work functions.
This is the important strategic distinction: memory does not have to mean vendor memory. Teams that deliberately keep knowledge and history portable can benefit from better models without starting over.
The question to ask after this week's episode
The podcast also covers Opus 5, Claude's incidents, Gemini Robotics ER 2, Grok Voice, Mistral Vibe, OpenAI Transcribe, GPT-5.6 pricing, Perplexity Model Council, and the week's important SDK and MCP changes.
One question connects them better than another model ranking:
If we changed models tomorrow, what knowledge about our work would we lose?
The answer shows where lock-in is already being built. If this resembles your situation, start by marking what is company knowledge and what merely happens to sit inside a vendor's project memory. That is useful input for Tool Forge when an AI workflow needs to become practical without trapping the knowledge.
This episode is an AI-generated masterclass built from Hammer's deep daily research into AI-provider updates and features, processed with NotebookLM.
Sources: Hammer's AI-generated masterclass research, based on daily deep research into each provider's updates and features.
FAQ
What does work memory mean in an AI platform?
It is the persistent context built around the model: project files, instructions, decisions, execution history, connectors, and sometimes agent identities. Its value is that work can continue without the team starting over.
Are AI models becoming easier to replace?
Yes. Claude Opus 5 appearing through several platforms, Perplexity Gateway API, and the stateless MCP architecture all point toward greater model portability. That does not mean project memory, history, and connectors are equally easy to move.
What should a team ask before choosing an AI platform?
Ask: If we changed models tomorrow, what knowledge about our work would we lose? Then check whether files, instructions, history, and permission logic can be exported in useful formats.
The Forge newsletter
Get new articles in your inbox
Pick the topics you care about. No noise, at most one email a week.
We follow GDPR. Unsubscribe anytime.


