AI Enablement Radar week 41: measure completed cases, not just AI usage

AI usage is a weak measure if nobody knows which cases were completed. Week 41 brings a concrete example: GitHub is fixing agent activity measurements that were partly missing or assigned to the wrong category. Meanwhile, Zapier shows how AI can classify incoming feedback and route it onwards. Our recommendation this week is to follow one case through to completion, including waiting and corrections, before judging the value.
Top signals this week
- On October 6, GitHub disclosed missing agent activity in Copilot metrics and some activity incorrectly counted as CLI usage. VS Code 1.139.0 includes the fix. Historical data cannot be repaired; billing was unaffected. Source: GitHub: restore agent activity metrics.
- On October 6, Cursor introduced remote control of local agents from its iOS app. The computer must remain on and connected. This is mobile supervision of local work, not a move to the cloud. Source: Cursor: remote control for local agents.
- On October 7, GitHub made local sandboxing generally available in Copilot CLI, the Copilot app and supported VS Code sessions using Agent Host. A sandbox restricts what a running agent can access, such as files and networks. Source: GitHub: local sandboxing.
- Zapier's October 6 guide describes using Gemini to classify feedback and send flagged cases to Slack. This is current integration guidance, not evidence that the entire capability launched that day. Source: Zapier: Gemini automation.
What organizations are actually doing with AI
Zapier trains the organization, not just the model
Zapier describes internal hackathons, training resources and leadership participation in its October 8 guide. The company reports 97 percent AI adoption, measured through recurring employee surveys. That tells us something about usage, but does not demonstrate profitability. For a team adopting AI, the useful question is which recurring tasks employees can now complete better.
Source: Zapier: AI adoption.
Live Fit Gym shows the groundwork automation needs
Jotform's customer story, updated October 9, describes Live Fit Gym expanding from one liability waiver to 60 forms and 14 tables across nine locations. The customer estimates total savings of 10–20 hours a week. This is workflow automation, not a reported AI outcome. The distinction matters: a well-organized intake process can be more useful than another model. Once information reaches the right person, it becomes easier to see where AI is needed.
Source: Jotform: Live Fit Gym.
The tooling layer: platforms, agents, and workflows
Let Gemini suggest a category and automation choose the recipient
Zapier's guide distinguishes AI by Zapier from a connector inside Gemini Enterprise. The latter requires the right license and administrator configuration; ordinary Workspace access is not equivalent. Start in the tool you already use. Let AI suggest a category and explanation for an incoming case, then use explicit rules to choose the recipient. An agentic workflow lets AI choose steps and use tools as well. Not every classification task needs that autonomy.
Source: Zapier: Gemini automation.
Count completions, not just agent activity
GitHub's measurement error makes historical comparisons unreliable for affected sessions. Update the development environment and mark the break in measurement before comparing periods. Then add your own operational measure: cases accepted by the right recipient without rework, for example. More recorded activity after the fix does not necessarily mean more people have started using AI.
Source: GitHub: restore agent activity metrics.
GitHub also introduced code review billing options on October 8. Organizations can pay rather than consume a member's entitlement, but this requires AI Credits paid usage. Member billing remains the default. Decide who owns the cost before making review a shared routine.
Source: GitHub: code review billing.
Mobile supervision still needs a working computer
Cursor lets users read and message local agents after pairing approved in the desktop app. Enterprise administrators must enable the feature. Plan who takes over when the computer is offline rather than treating the mobile app as an independent operating environment.
Source: Cursor: remote control for local agents.
Governance and risk: what needs to be in place before scaling
GitHub's local sandboxing provides technical boundaries for agent access. Organizations can require policies developers cannot weaken. Pair this with scoped permissions, keys in environment variables or a secret manager, and a run log that redacts sensitive values. Human approval belongs at customer commitments and other decisions the workflow has not been authorized to make.
Source: GitHub: local sandboxing.
On October 9, the European Commission published information about a special meeting of its scientific panel on frontier AI. The panel has investigated reported loss-of-control incidents. The announcement does not contain new obligations or completed recommendations. Our practical takeaway for buyers is to agree how the supplier will report incidents and model changes.
Source: European Commission: scientific panel meeting.
Two older guidance documents are useful here. The Commission's AI standards page distinguishes voluntary standards from harmonized standards that, following the required process, can provide a presumption of conformity. UK procurement guidance from 2020 recommends logs, monitoring model performance and supplier knowledge transfer. Neither is a new rule from this week.
Source: European Commission: AI Act standardisation.
Source: UK Government: AI procurement guidelines.
This week's practical Hammer test
Set aside 40 minutes to test the routing of incoming questions. Choose course inquiries, service cases or quotation requests, for example. The aim is to discover where work gets stuck, not to demonstrate the most AI features.
- Spend the first 10 minutes choosing five completed cases you are authorized to use. Record the correct recipient and what counted as completion.
- Spend 10 minutes asking AI to suggest a category, the next responsible role and missing information. Work in a draft without external messages.
- Spend 15 minutes comparing the suggestions with the actual decisions. Count correct recipients, necessary corrections and time to a usable suggestion. Have a colleague review uncertain cases.
- Use the final five minutes to assign an owner for a week-long test. Decide which cases automation may route and which require approval.
Copy this instruction into your approved AI tool:
Read our categories and the five cases. Suggest a category and the next responsible role for each case, supported by the material. Flag missing information and uncertain cases. Draft a tracking list with recipient, correction and completion status. Send nothing and leave the originals unchanged.
If the team already has an integration, save suggestions in a separate review view. Use limited access and log the decision when a person approves it. After a week, compare rework and time to completion with your previous routine.
Companies and tools to watch
- GitHub: check how the break in measurement affects your reporting. Source: GitHub: restore agent activity metrics.
- Zapier: try one explicit classification step in an existing workflow. Source: Zapier: Gemini automation.
- Cursor: assess whether mobile supervision reduces waiting for human decisions. Source: Cursor: remote control for local agents.
- Jotform: use Live Fit Gym as an example of organized intake before more advanced AI. Source: Jotform: Live Fit Gym.
Need to connect the test to your forms, cases and responsible colleagues? Through Tool Forge, Hammer Automation helps build a workflow that shows what gets completed and what still needs a person.
The Forge newsletter
Get new articles in your inbox
Pick the topics you care about. No noise, at most one email a week.
We follow GDPR. Unsubscribe anytime.


