AI Enablement Radar week 37: testing AI for consistent quality

AI Enablement Radar week 37: testing AI for consistent quality

AI is becoming easier to reach while work is happening: Gemini is arriving on Windows and in Android spreadsheets, while Cursor is starting to coordinate agents at project level. For teams that have already tried AI, the next step is to see whether a task works equally well on repeated runs. This Radar connects the September 7–13 updates to concrete examples of training, fact gathering, and repeatable quality checks.

Sources: Google on Gemini for Windows. Google on spreadsheets on Android. Cursor Projects.

Top signals this week

  • Cursor Projects began rolling out in beta on September 10. A coordinator plans and delegates; other agents do the work. Shared context files connect local and cloud agents. This suits recurring project tasks, but the beta status belongs in your rollout plan. Source: Cursor Projects.
  • GitHub added usage metrics for the dedicated VS Code Agents window on September 11. Session and message counts show activity, not quality or time saved. Source: GitHub's new usage metrics.
  • Notion gave Business and Enterprise workspace owners control over models available to Notion Agent and Custom Agents on September 9, plus a default model for Custom Agents. Source: Notion model controls.
  • Gemini for Windows, announced on September 11, opens with Alt + Space. Google describes project summaries using Gmail and Drive material. This is a new way to reach the assistant, not an announcement of unrestricted autonomous computer control. Source: Google on the Windows app.
  • Gemini in Google Sheets on Android can answer data questions and create charts. Complex edits, formatting, and new formulas remain web tasks. Rollout began on September 9 and may take up to 15 days. Source: Google on Sheets on Android.

What organizations are actually doing with AI

The following customer stories appeared before this week. They are current comparison examples, not new launches. Their figures come from vendor customer interviews rather than independent productivity studies.

Pictet counts training separately from access

In Anthropic's September 2 customer story, around 700 people at Pictet have access to Claude Code or Cowork. More than 500 have been trained through 25 hands-on workshops with Artefact. Keep the distinction between accounts and training in your own reporting. Then add how many people actually use the routine and whether the output holds up.

For a team, the equivalent could be a shared working session where everyone reviews the same proposal material. Differences in instructions and judgment become visible without buying more licenses to learn something useful.

Source: Anthropic on Pictet's rollout.

Spellbook tests the same contract repeatedly

Spellbook reports reviewing around 530,000 contracts a month. Its method is more useful to someone planning a test: the company runs the same contract and instruction ten times to check whether the system detects issues consistently. It also tracks which suggestions users accept.

A well-written answer is not enough of an evaluation. Ask whether AI finds the same important discrepancy when the task is repeated. The customer story was published on August 27.

Source: Anthropic on Spellbook's contract reviews.

EvenUp gathers facts before writing

EvenUp describes a pipeline that reads documents, classifies them, reconciles duplicate facts, and keeps source-page references before producing a draft. Attorneys review and sign every draft. The August 26 story offers a useful pattern for other document workflows: build a checked fact set and reuse it in the summary, customer letter, and internal follow-up.

Source: Anthropic on EvenUp's document workflow.

The tooling layer: platforms, agents, and workflows

An agentic workflow lets AI choose and perform steps through tools rather than only writing an answer. Cursor Projects is this week's clearest example of longer coordination: projects can respond to schedules, Slack, and pull-request activity. Our recommendation is to start with a recurring deliverable, such as a weekly list of documents that need updating. Define what a finished delivery should contain before setting a schedule.

Source: Cursor on projects and shared context.

Google is making everyday questions easier to ask. A field manager can review a spreadsheet on Android; a colleague on Windows can assemble project material. Check the account, subscription, and actual availability before a team test. Mobile support does not replace every web feature, and the Windows app follows the organization's existing generative AI controls.

Sources: Google on Sheets on Android. Google on Gemini for Windows.

GitHub's new statistics need careful interpretation. They cover the dedicated Agents window, not Agent Mode in the regular editor window. Missing fields may be absent or null and should not automatically count as zero usage. Pair activity metrics with a separate check of completed work.

Source: GitHub on the scope of the new metrics.

Governance and risk: what needs to be in place before scaling

AI governance here means who chooses tools, sets access, and owns the result. Notion's model controls make one of those choices shared. GitHub's September 9 update lets Business and Enterprise administrators block, allow, or require approval for operations including commands, file actions, and network domains in supported agent environments. Check that your client is covered; a central setting only helps where it is enforced.

Sources: Notion model controls. GitHub managed permissions.

For schools, Skolverket's advice, updated on August 31, provides relevant background: keep written, revisable guidelines, and do not let AI replace teachers' professional judgment or grading decisions. These are recommendations, not new rules introduced this week.

Source: Skolverket's advice on AI.

For an integrated test, we recommend scoped permissions, keys in environment variables or a secret manager, and redaction of fields the task does not need. Keep a short run log. A named person approves external messages and other binding actions.

This week's practical Hammer test

Repeat the same document task in 40 minutes

Choose a recurring task, such as comparing a proposal against an approved fact sheet. The aim is to uncover variation, not prove a general time-saving claim.

  1. 0–8 minutes: Choose the material and write down the facts and discrepancies a useful answer must capture. Have a knowledgeable colleague establish the reference answer.
  2. 8–20 minutes: Run the same instruction three times in fresh conversations with the same model and material. Ask for source references. Do not change the instruction between runs.
  3. 20–32 minutes: Compare the outputs against the reference answer. Note missed facts, invented claims, and how much correction each response needs. Three runs are an initial check, not statistical proof of reliability.
  4. 32–40 minutes: Improve the part that varied most: the source material, instruction, or review step. Save the version and schedule another test on different documents. Keep everything as drafts during the comparison.

A short instruction to adapt:

Read the material and compare it with our review criteria. Ask first if criteria or sources are missing. Cite the source for each discrepancy and distinguish supported facts from points you cannot resolve. Produce a review draft without changing source files or sending anything.

Companies and tools to watch

  • Cursor Projects: watch whether project context holds up as multiple agents take on different parts of the work.
  • Pictet: borrow the distinction between access to the tool and practical training.
  • Spellbook: use repeated runs as a model for your own quality checks.
  • Google Gemini: test whether Windows and Android access removes a real interruption in the working day.

If the test finds recurring differences, Hammer can help improve the instruction and review method through Skill Forge. The next step is a routine a colleague can use and assess, rather than just a better answer in your own chat.

The Forge newsletter

Get new articles in your inbox

Pick the topics you care about. No noise, at most one email a week.

Get new articles in your inbox

We follow GDPR. Unsubscribe anytime.