
Cutting AI Costs Without Cutting Output For Amazon Sellers
Danny McMillan and Sim on model switching, OpenRouter cost hacks and whether AI spend for Amazon teams is actually worth it.
A quick solo show this week — Dorian's out sick, Matt's unavailable, so it's just Danny and Sim covering how to control AI spend without gutting output.
Sim walks through the real cost pressure of a 25-person team on Claude: five figures a year, uncapped Fable usage, and no way to see per-person burn. Danny counters with the setup he's built to fix exactly that — routing grunt work through OpenRouter and Kimi K3 without preloading tools, which took one job from £4-5 down to 35p. They land on a practical split: Fable 5 for planning, a cheaper model for the build, staged "cascade" plans to keep context windows under control. Sim also shares a genuinely wild same-day case study — a 100-video yoga app build for about £300 — and Danny pushes back on the whole framing: stop looking at AI as a cost line and start looking at what it's generating.
Key Topics
Picking one AI provider and sticking with it - the cost of switching a trained team, and why Astra being better than Fable isn't reason enough to move
Why teams actually burn tokens - heavy browser automation, parallel video work, and vibe-coders scaling far past their job description
The OpenRouter cost hack - Kimi K3 without preloaded tools, cutting one job's cost by roughly 90%
Cascade planning - staging big builds across multiple sessions to control context window burn
Chinese model options - GLM 5.3 and LM Studio AI as free-to-cheap alternatives
Grok as a low-friction accessory - personal automation for AI-hesitant team members
ROI framing - weighing AI spend against what it actually produces, not just what it costs
Timestamps
[00:01] Quick solo show - Dorian ill, Matt unavailable, flight to Mallorca later
[01:02] Wishing Dorian a speedy recovery, back next month
[01:17] Kicking off: token costs, Astra's release
[01:42] Why constant provider-switching burns team trust and training
[02:10] History: Google/Gemini 18 months ago, Claude masterclass in December, full team switch
[02:29] Sim's Opus 5 complaint - unclear responses, had to cross-check with Sonnet
[02:58] Fable's cost problem - no way to cap runaway usage
[03:28] Costs into five figures a year; splitting power users vs Chinese-model users
[03:54] The NVIDIA DGX Spark purchase for local model runs
[04:32] Danny's setup: Kimi in the terminal, Claude Desktop for daily work
[04:59] The two real causes of heavy token burn: browser automation, parallel video editing
[05:58] The "I've got ADHD" repo - controlling how Claude communicates
[06:25] Claude Desktop's concise-response setting, plus widgets vs terminal mind-mapping
[07:27] Auditing what each team member actually uses AI for before cutting costs
[07:56] Why max-plan tool loading is free but third-party API tool loading isn't
[08:26] "OpenClaude" - a terminal alias that loads Kimi without preloading tools
[09:24] Sim's example: a warehouse dispatcher now vibe-coding time-tracking apps
[09:54] The $20-to-Max "purgatory" middle ground
[10:44] Danny's recommendation: Kimi K3 for the team, Claude Desktop kept for flexibility
[11:12] Splitting by task: Fable 5 for planning, K3 for the build
[11:40] Cascade plans - staging big jobs across sessions to control context
[13:04] Real numbers: four parallel video edits, £4-5 down to 35p after removing tool preload
[14:07] LM Studio AI and the free-credit-then-continue usage hack
[14:43] GLM 5.3 - "Opus 5 good," and it speaks plain English
[15:14] Grok/Grokbot - zero-friction agent automation from a phone
[16:21] Building a review-flagging SOP from a LinkedIn post in one sitting
[16:38] API scraper tip: Apify or Monid for cheap review scraping
[19:00] Team breakdown: 5 on Max plans, ~18 on $20 plans
[19:22] The frustration of no pooled token allocation across a team
[20:12] The "Claude effect" - wait 21 days before reacting to a new release
[21:08] Astra vs Fable, and why a full company switch is a bigger decision than it looks
[22:04] Reframing: only 3-4 people actually need heavy-lifting models
[23:37] Danny's pushback: what is the AI spend actually generating?
[23:59] Case study: ~£300 to edit 100 yoga app videos with real footage and AI voiceover in a day
[26:13] ROI framing: AI cost against revenue and headcount value
[26:45] Team-wide Max plan cost: roughly £30k/year - "just a salary"
[27:08] The local GPU reality check - a 90GB model locks the RAM, breaks other flows
[28:03] RunPod and HyperStack as pay-by-hour alternatives to token metering
[28:34] Wrap-up - Danny off to the airport
Key Takeaways
Constant model-switching costs more than it saves - retraining a team and rebuilding skills outweighs chasing the newest release; wait it out, then decide.
Tool preloading is where the real API cost hides - stripping it out cut one job's cost by roughly 90% with no drop in output.
Split the model by the task, not the team - planning on a stronger model, execution on a cheaper one, staged across sessions to control context burn.
Audit usage before cutting spend - most token burn traces back to a handful of people doing genuinely heavy work, not the whole team.
Cost only means something next to output - a £300 same-day build replacing two weeks of manual work reframes what "expensive" actually means.
Notable Quotes
"The last thing you want to be doing is getting people affiliated with the software, get them used to it, and then suddenly taking that away." - Sim
"Once you took that out of the equation, I've got it down to about 35 pence for the same work." - Danny McMillan
"Have you looked at what it's adding?" - Danny McMillan
"It's not even remotely comparable." - Sim
Resources Mentioned
OpenRouter - model-switching gateway, used here to route to Kimi K3 without preloaded tools
Kimi K3 - the cost-efficient model for day-to-day build work
GLM 5.3 - Chinese model tried via LM Studio AI, described as "Opus 5 good"
LM Studio AI - harness for running Chinese models, including a free-credit usage workaround
Grok / Grokbot - low-friction personal AI agent with built-in scheduling
Apify / Monid - API scrapers used for cheap-at-scale review data pulls
NVIDIA DGX Spark - local hardware for running in-house model flows
RunPod / HyperStack - pay-by-hour GPU rental as an alternative to token-based pricing
Connect
Sim - Amazon seller and co-host
Seller Sessions is the leading podcast for advanced Amazon sellers, hosted by Danny McMillan.
Mais episódios de "Seller Sessions Amazon FBA and Private Label"



Não percas um episódio de “Seller Sessions Amazon FBA and Private Label” e subscrevê-lo na aplicação GetPodcast.








