AI Masterclass Podcast Week 30: habit may decide the next AI winner

AI Masterclass Podcast Week 30: habit may decide the next AI winner

A benchmark can tell you which model won a test. It cannot tell you which assistant you will actually use at 8:17 in the morning, when the forecast is open, the meeting has just ended and the work already lives in Excel.

Week 30 points to a different kind of AI competition. Google is moving Gemini deeper into Samsung devices. Grok is appearing in Excel and Outlook. Claude is connecting its voice mode to tools. OpenAI is building health and enterprise workflows around ChatGPT. Perplexity is making reusable skills part of its Agent API. Manus lets one agent work across multiple Google accounts.

This week's thesis is that habit may decide the next AI winner more than benchmark scores do. If that is right, the important question is no longer only “which model is smartest?” It is also “which service is already in the right place when the work needs to get done?”

The AI battle is moving into existing habits

The clearest signal comes from Google. In the company's own reporting, the Gemini app had 950 million monthly active users, while more than 40 apps are part of Samsung's expanding AI surface. Those are vendor-reported figures, not an independent measure of value. They still show the advantage of putting AI inside the phone, documents, and operating system.

The same movement is visible at work. Grok has launched Excel and Outlook add-ins for paying X and SuperGrok users. Anthropic has expanded Claude's voice mode with tool connections. OpenAI Health brings health-related context into ChatGPT for adult users in the United States, while the enterprise product Presence packages agents, handoffs, and implementation support in a limited release.

This is not one uniform market. It is several attempts to win the same moment: the moment when a user already has the file, inbox, calendar or phone open.

Placement gives a model the first chance

A better model can still win on quality. But a model that requires people to leave their workflow must first win another contest: it has to be remembered, opened and given the right context.

That is why distribution is starting to look like a moat. Mistral is strengthening its enterprise route through Microsoft Foundry and Copilot Studio. OpenAI is using partners such as KPMG and a more service-heavy product in Presence. Google combines Android, Samsung, Workspace and its own models. Manus takes another route, turning the agent into a router across as many as ten Google Workspace accounts.

If the thesis holds, buying decisions will depend less on a single leaderboard position and more on three ordinary questions:

  • Where does the work begin? In a document, email, conversation or business system?
  • Does the context travel with it? Or does someone have to rebuild the task in a new chat every time?
  • Can the routine move? A habit that works with only one vendor can become expensive when prices, models or terms change.

Reusable routines are the answer to lock-in

Perplexity's Agent API shows the other half of the competition. The API can use skills that describe how a repeated job should be done and create Office files inside a durable workspace. At the same time, Perplexity is steering developers away from the older Sonar interface and toward the new agent platform.

That is convenient, but convenience and lock-in often grow together. When instructions, files, and run history become part of a platform, switching costs more than changing a model name.

The practical response is not a large control exercise. It is to separate the routine from the engine. Describe what material goes in, what a completed result looks like and which parts must survive a vendor change. The same way of working can then be tested in ChatGPT, Gemini, Claude, Grok, Perplexity or Manus without starting from zero.

Mass adoption is not the same as deep work

There is an obvious weakness in this week's thesis. Google's ATLAS study covered roughly 15 million interactions and found broad use across occupations, but fewer than one in ten interactions attempted to automate a task from start to finish. Most people were still using AI as a collaborator.

A place in the phone or office suite therefore does not automatically create a good workflow. Raw model capability still matters. A model can also lose features when it is sold through another platform; Perplexity, for example, warns that not every native Gemini capability is guaranteed through Agent API.

Habit may decide who gets the first attempt. Quality, cost, and the ability to finish the work decide who gets to stay.

A better test than another model ranking

Choose a repeated task that already happens in a tool you use: summarize a customer thread, compare forecast with actuals or prepare the first draft of a meeting brief. Run it where the work already lives and compare it with your current process.

Measure two things: time to a reviewable result and how much has to be redone. Keep the task description outside the vendor's chat. If the test works, you have learned something about both value and switching cost.

If your work keeps getting stuck between several tools, Hammer's Tool Forge can help turn it into a clear, practical routine without starting with a large platform investment.

This episode is an AI-generated masterclass built from Hammer's deep daily research into AI-provider updates and features, processed with NotebookLM. The article complements the podcast and uses the source material plus a separate NotebookLM synthesis; no audio transcription was used.

FAQ

What is the main thesis of the AI Masterclass Podcast for week 30?

Distribution and habit may become a stronger competitive advantage than a single top benchmark result. The service already present where work happens often gets the first chance.

Does that mean model quality no longer matters?

No. Habit may decide which service gets tried first, but quality, cost, and the ability to complete the task decide whether it stays in the workflow.

What can a smaller organization test after this episode?

Choose one repeated task inside a tool you already use. Compare time to a reviewable result and how much must be redone. Keep the task description outside the vendor’s chat so the routine can move.

The Forge newsletter

Get new articles in your inbox

Pick the topics you care about. No noise, at most one email a week.

Get new articles in your inbox

We follow GDPR. Unsubscribe anytime.