I used to work as a finance consultant, years in Excel where one wrong number costs real money. So I didn't ask Claude or ChatGPT to write an email. I gave both the exact same four Excel jobs a real business deals with: messy data, a broken formula, a dashboard, and a financial model with hidden mistakes. Same files, same prompts, scored out of five per task. ChatGPT ran on the latest 5.5 model with extended thinking; Claude on Opus 4.8 with matching thinking effort.
Round 1: Cleaning a messy sales export
220 rows with dates in four different formats, uncapitalized customer names with stray spaces, inconsistent country names, quantities stored as text, prices with mixed comma/dot notation, and missing revenue formulas.
Result: a tie. Both fixed all six problems perfectly, and both independently found and removed the same 12 duplicate rows nobody asked about. ChatGPT was faster (3 min vs 5 min).
Round 2: Fixing broken formulas
A receivables workbook with a #VALUE! error in the days-overdue column, a wrong total outstanding calculation (ignoring partial payments), and a summary table needing a SUMIFS across multiple conditions for invoices more than 30 days overdue.
Result: another tie. Both nailed every fix, and, remarkably, both wrote the same IFERROR + INDEX/MATCH and SUMIFS formulas. Different AIs converging on identical Excel formulas.
Round 3: A management dashboard
Clean transaction data; the task was a one-page dashboard, revenue, gross profit, margin, revenue by region and product, top 3 sales reps, monthly trend, at least one chart, everything as live formulas.
Result: Claude edges ahead. Both were correct and fully dynamic, but Claude produced KPI boxes, removed gridlines and drew three charts instead of one confusing rainbow-colored chart. The lead was purely on design.
Round 4: Auditing a real 5-year real-estate model
My own model built for a client, construction plus selling phases, monthly over five years, revolver financing, 150+ rows of formulas, with three planted mistakes: an accumulated payment double-counting a building, a first-year interest formula not divided by 12, and maintenance costs missing the inflation factor. The prompt just asked to check before sending to the client.
Result: Claude found all three. ChatGPT found one, and hallucinated errors that weren't in the model at all. This was the round that mattered.
Verdict
As of June 2026, Claude wins for serious Excel work, especially anything sophisticated where a hidden mistake costs money. ChatGPT consistently beat Claude on speed for the easy tasks, so it remains a fine choice for quick-and-simple jobs.
One takeaway regardless of tool: AI states a wrong number with the same confidence as a right one. Let it do the heavy lifting, then check the numbers before they leave your desk. That habit is the difference between saving hours and shipping a mistake.
