01
Check
The Claude skills or plugin you already use, checked.
- Validation and a linter for silent failures
- A routing test on the phrases your team uses
- A short report: what misfires, and why
Claude skills
Claude is good at arithmetic; what goes wrong is what it counts: test orders, VAT, which month a refund or a payout belongs to. I build Claude skills that compute every figure with a script, name every order left out and print your definitions under every answer. Or I check the skills or plugin you already use.
Three synthetic Shopify stores, nine owner questions, three runs each with Opus 5.5, with the skills and without. Every answer is public.
Fixed price, quoted per export.
01
The Claude skills or plugin you already use, checked.
02
One month-end skill built on your export.
03
All three skills, with every check.
The question your team asks Claude every month, and a sample export. Test data or an anonymised copy is fine.
You see what was counted and what was left out before anything is built. The skills have not yet run on a real store's raw export, so this step comes first.
Three runs per question, mistakes planted on purpose, every answer read.
A plugin for Claude Code, its tests, every answer and grade, and install notes.
Synthetic stores, Opus 5.5 only, three runs per question: a repeated behaviour, not an error rate. The same person wrote the skills and the tests; every answer and grade is public, so anyone can re-grade them. Your files are read by Claude for your order only and deleted after delivery unless you ask me to keep them.