👋 Greetings {{first_name|earthlings}},

In a break to normal proceedings, this edition has been hijacked by Asta and Chris

We are both part of Spark AI’s internal team. Asta is in marketing, using AI mostly for writing and research. Chris is in automations and operations, using it to build internal processes and systems. We thought that made us a reasonable pair to answer something that we often get asked by our clients at Spark: What AI model should I be using? 

While we’ve mainly been using Claude internally, we spent the past two weeks properly testing the latest in ChatGPT and Gemini, noting what we found and each one’s strengths. We haven’t tested Copilot in this head-to-head because we’re in a Google environment so most of its integrations don’t work for us. But if you’ve got a Copilot question, do let us know - because we use all four platforms across our programmes with our clients. 

Here's what we’ll cover in today's issue: 

Why the ‘best’ model matters less than where you build your context

Our answer to the ‘which AI platform is best’ question is always the same. Usually someone is deciding which Large Language Model (LLM) their agency should be using. Our mantra is simple: the best one is the one you are already using, the model already in your tech stack.

It’s easy to get distracted by the cool new thing (Claude Cowork! ChatGPT Agents!) but they are all constantly evolving and playing catch-up with each other. Less than two weeks ago, Claude overtook ChatGPT in US business AI spending for the first time. Then last week Gemini dropped lots of new features. So the rankings shift all the time and it's easy to get FOMO. What matters more than which model is ‘best’ right now. is how deeply you have built your working practice and your internal systems inside one of them.

Every time you switch tools, you lose context and you lose fluency. There is real value in knowing how a model operates day to day and how to navigate its features. That capability and confidence using AI shows up in the outputs. Saying that, we did not come back from two weeks of testing with nothing to report. We thought it would be useful, for us and for you, to take a look at how the models compare right now.

So here is what we found. 

Claude

Claude produces the strongest writing quality of the three. Sentence variation, tone, and flow feel more human than the other models. It also produces solid final versions of documents. The skills feature is particularly useful for embedding consistent brand knowledge or ways of working, without re-explaining it each time. The MCP (app connectors) setup is also very clean here. You can control which tools are active in each chat, and what permissions Claude has, which helps keep you in control.

The main practical constraint is token use. On the Team plan, Conversations and Claude Cowork/Code consume tokens quickly, and once the limit is reached, you have to wait for them to reset before you can continue or pay more. The way around this is to upgrade each team member from a Standard to a Premium subscription, but then you’re looking at £70 per month per user. Despite us noting that Claude is good at producing final outputs, that strength can make it feel less collaborative. It tends to push toward creating a finished document a bit too soon rather than working ideas through with you.

ChatGPT

ChatGPT makes the process of getting to a good output feel straightforward. The new Agent feature is a big step forwards from custom GPTs, and building agents them is now guided and intuitive. ChatGPT now also supports skills -using the same format as Claude - so you can quickly train it to perform tasks the way you want it to. Meanwhile Codex brings coding capability into reach for non-developers. Where ChatGPT really excels is in the middle of the process. It feels more like a collaborator than a delivery engine, better suited to shaping ideas, thinking things through, and iterating before you know what the final output needs to be 

However, we found that final outputs are weaker. The writing quality does not quite reach the same standard, and producing a polished final document takes a bit more work. The MCP connector settings (where you connect up to the rest of your tech stack) are also more fiddly to get to, and harder to configure, which makes life that bit more difficult.

Gemini

Gemini's biggest strength is how embedded it is in everyday work. It shows up across your Google Workspace, from search and gmail to meeting notes and slides. We rely on these integrations a lot, and we are impressed by its outcomes. If your team already works inside Google, Gemini reduces friction without requiring a change in habit.

Gemini’s killer feature is NotebookLM. It’s worth getting a Workspace account just for this. An AI notebook that can hold 300 documents with all the answers grounded and referenced back to the source. We use one for every client, and you probably should too.

But as a standalone tool, Gemini has more limitations. It is weaker at following instructions consistently, and getting to a good output requires much more iteration. However its answers are much more to the point than ChatGPT - so you’re less likely to drown in information. Google held their annual I/O conference last week and interestingly, several of the weaknesses we identified during our research around lack of personalisation, agents, and accessible coding capability, are being addressed. Scroll down to read our tool updates to find out what they have implemented and what it might mean in practice.

All the features and what they do

Across these three LLMs, there are a lot of features and knowing which one to reach for is half the battle. We put together a summary for each model that breaks down what each feature does, how to think about it, and a practical example of when to use it. 

Including a table of each model would make this newsletter even longer than usual so we’ve put them all on a dedicated webpage for you.

Audit your LLM in 15 minutes

So, this week's task is simple. Take 15 minutes to audit the tool you rely on most and strengthen what it knows about you. It is a small investment that will save you a lot of future iterations. 

Open up your chosen LLM and find a tool you use often, whether that's a Project, a Custom GPT, a Skill, a Gem, etc. 

Ask it these questions to improve its instructions and the context it is working with:

  1. What is missing from my context that would strengthen the reliability of your outputs?

  2. What assumptions are you making about the way I work that could be made explicit in your instructions?

  3. What do I always correct you on mid-conversation that could be built into your instructions so you get it right the first time?

Then add context yourself:

  1. Upload or paste an example of an output you were happy with and explain why it worked.

  2. Share a writing guide, a brand document, or a style reference so the model produces work with brand consistency. 

What we’ve been up to

New podcast episode out now

Last week, Our founders Jules and Emma  dropped the latest episode of their podcast  What’s New in Creative AI. 

The discussed the paradox that almost 90% of agency staff are now saving time with AI. But 35% still say lack of time is what's holding them back from going further. Think that's weird? Tune in to find out more. Alongside that they gave advice on how to quick-start your AI governance and what AI capability really looks like.

As always, it’s packed with practical guidance on how agency leaders can turn this AI activity into a competitive advantage.

Taking bookings for Cannes Lions

Emma and Jules are heading to Cannes Lions this June. They are opening up 1:1 free AI audit sessions for agency leaders who want a clearer picture of where their business really stands with AI and what to do next.

Using our Spark AI Maturity Model, alongside our own custom built AI audit tool, Emma and Jules will walk through your agency’s current level of AI adoption with you, where the gaps are, and what needs to happen next to turn AI use into a competitive advantage.

All the tool updates

So far, we have dedicated this newsletter to all things tools, so it is only right to keep that momentum going by sharing the latest AI updates. 

Google

Google held I/O, it’s annual developer conference, last week and announced what Sundar Pichai called "the agentic Gemini era." Strip the framing away and what they actually announced was a roadmap to catch up with ChatGPT and Claude, not a product you can use next week. For UK agencies on Workspace Business, the gap between the keynote and your inbox is wider than the marketing suggests. Here's the honest translation:

Gemini Spark. This is Google's answer to Claude Cowork and ChatGPT Agent mode: a cloud-based agent that runs 24/7, takes a goal and executes multi-step jobs in the background, even while your laptop is shut. It connects natively to Gmail, Drive, Docs, Sheets, Slides, Calendar, YouTube and Maps. It introduces "Skills" as composable, reusable instructions, which is the same vocabulary Anthropic uses, and a clear sign of where the industry has landed.

The catch: Spark is currently in trusted-tester beta, with a wider rollout to US-only Google AI Ultra subscribers ($100/month). For Workspace Business customers in the UK, Google's language is "available soon in preview for business customers this summer." No firm date. ChatGPT and Claude have had their equivalent agentic products generally available for months. Google has announced parity but hasn't shipped it.

The genuinely interesting model for creatives is Gemini Omni. Unlike Veo, which was text-to-video, Omni is a multimodal model that generates editable video from any combination of text, images, audio, and existing clips. The practical difference: every edit instruction is understood in context, so you can change the camera angle on step three and it knows the characters and lighting from steps one and two. No "start over." Currently capped at 10 seconds, watermarked with SynthID, and now live inside Google Flow alongside Whisk and Imagen. Omni Flash is rolling out globally to AI Plus, Pro, and Ultra subscribers, plus free on YouTube Shorts and Create. Your can get hands on this week.

What's actually landing on Workspace Business? Mostly previews and promises. Google Pics (object-level image editing inside Docs and Slides), voice features in Gmail, Docs and Keep, and a smarter AI Inbox are all "preview for business customers this summer." Gemini 3.5 Flash is the new default model behind your existing tools, which gives you a quiet quality upgrade without changing anything. The macOS Gemini app is downloadable today for everyone, but doesn’t give you access to Gems so we don’t recommend it.

The agentic infrastructure, the Workspace Intelligence admin controls, the Antigravity Agent Platform: all of that sits behind a Gemini Enterprise licence, which is the £21/user upsell sitting on top of your existing Workspace bill.

What this means for you: Google's deepest advantage is data integration - Spark touches Gmail, Drive and Docs natively, no connectors needed. But the deepest disadvantage is the speed at which it’s rolling these features out. By the time Spark reaches UK Workspace Business, the other two will have shipped their next iteration.

Anthropic 

Two announcements from Anthropic this fortnight, and together they make the positioning clear. Anthropic is building Claude as the AI for business, at every scale.

At the enterprise end, Anthropic announced a global alliance with KPMG, one of the world's largest professional services firms. Claude is being embedded into KPMG's central client delivery platform and rolled out across the entire global workforce, starting with tax, legal, and private equity work. This is AI embedded into the core of how a major firm delivers client work, not a side tool sitting next to it.

At the other end, Anthropic launched Claude for Small Business, with integrations into QuickBooks, PayPal, HubSpot, Canva, Google Workspace, and Microsoft 365. This is the kind of operational AI a lot of our clients are looking for. Anthropic is making it easier to get started and signalling clearly that Claude is where businesses should be building.

OpenAI

ChatGPT has recently improved a great deal, with GPT 5.5, agents and skills. These bring it’s “get stuff done” abilities up to a similar level to Claude.

OpenAI also launched the Deployment Company, a new business focused on helping organisations embed AI into real workflows. It’s recognition that getting the value from these technologies is not straightforward - the very reason Spark exists!

They’ve also caught up with Claude Code by bringing Codex to mobile, so teams can review, steer, and approve coding work from their phone. It makes AI agents part of everyday working rhythms, including time away from a desk.

That’s it for now, see you next time!

Asta and Chris

Spark AI

PS if you’d like to find out more about our AI programmes for agencies and brands download our brochure.

Spark AI Brochure Spring 2026.pdf

Spark AI Brochure Spring 2026.pdf

2.41 MBPDF File

About Spark AI

Spark AI helps you lead your team through the biggest shift since digital, through AI training, transformation and workflows. We've worked with 70+ agencies, published the #1 bestselling book on AI for Agencies, and teach at Oxford University.

Not a subscriber yet?