Claude vs Gemini for engineering calculations: a hands-on test
TL;DR — I built six calculations designed to catch the mistakes engineers actually make: gauge versus absolute pressure, a forgotten static head, a normal-versus-actual volume conversion, vapour pressure at temperature. Claude Opus 5 and Gemini 3.1 Pro both went 6 for 6. Neither fell for a single trap. The useful difference was not accuracy — it was how much each one told me that I had not asked about, and what each did when I handed it a physically impossible specification.
I am a plant and mechanical engineer. Thirteen years on industrial machinery and energy systems, now running an independent practice where both of these models are open on my second monitor most of the day. These are the calculations I actually do, not a leaderboard benchmark.
Why the leaderboards don't answer this
Published benchmarks measure competition math and code. Neither predicts the thing that matters at a desk: whether a model carries units through a multi-step calculation, notices that a stated premise is physically impossible, and says so instead of producing a confident number.
An engineering answer that is 8% off is not "mostly right." It is a number that goes onto a data sheet, gets quoted to a vendor, and comes back as a purchase order. So the question I care about is narrower than "which model is smarter": can I hand this to a model and trust the output enough to check it in one minute instead of redoing it in twenty?
Test setup
| Models | Claude Opus 5 · Gemini 3.1 Pro |
| Date tested | 20 August 2026 |
| Interface | Web chat app, both |
| Settings | Claude with thinking effort set to high; Gemini at its default — see limitations below |
| Prompt style | Identical prompt text to both, all six problems in one continuous conversation per model, no follow-up coaching |
| Reference | Every answer checked against my own hand calculation and published steam tables |
Each problem was pasted in verbatim, in the order below, as one continuous session with each model — the way you would actually work through a set of calculations rather than the way you would run a clean benchmark. I did not tell either model it was being tested, and I did not correct an answer and re-ask. The first response is the response, because that is how it works when you are busy. That choice has consequences, and they show up in problem 6.
The six problems
Four of them carry a deliberate trap — the kind that produces a plausible-looking number rather than an obvious error.
- Saturated steam properties at 0.8 MPaG. Trap: gauge versus absolute. Reading the table at 0.8 MPa absolute gives 170.4 °C instead of 175.4 °C.
- Pipe pressure drop — 40 m³/h of 30 °C water, 250 m of ID 105.3 mm carbon steel, eight long-radius elbows, outlet 12 m above inlet. Trap: the 12 m static head, which dwarfs the friction term.
- Unit conversion — a burner rated 350,000 kcal/h, gas at 45 MJ/Nm³, meter running at 45 °C and 20 kPaG. Trap: normal versus actual volume, where the temperature and pressure corrections pull in opposite directions.
- Heat exchanger — counter-current, 12,000 kg/h cooled 95→60 °C against 18,000 kg/h of 32 °C cooling water. Duty, LMTD, area at U = 1,100 W/m²K.
- Pump shaft power and NPSHa — 25 m³/h of water at 80 °C, 45 m head, 72% efficiency. Trap: vapour pressure at 80 °C is 47.4 kPa, which eats half the atmospheric head.
- An impossible specification — condense 1,000 kg/h of steam at 0.5 MPaG and subcool the condensate to 28 °C, using cooling water available at 32 °C. A useful model refuses. A dangerous one calculates.
Results
| # | Problem | Claude Opus 5 | Gemini 3.1 Pro |
|---|---|---|---|
| 1 | Saturated steam | Correct — 175.4 °C, 2774 kJ/kg | Correct — 175.4 °C, 2773.9 kJ/kg |
| 2 | Pressure drop | Correct — 15.9 m, 155 kPa | Correct — 15.9 m, 155.5 kPa |
| 3 | Unit conversion | Correct — 407 kW, 32.6 Nm³/h, 31.7 m³/h | Correct — 406.8 kW, 32.5 Nm³/h, 31.7 m³/h |
| 4 | LMTD and area | Correct — 33.5 K, 13.3 m² | Correct — 33.5 K, 13.25 m² |
| 5 | Pump power, NPSHa | Correct — 4.14 kW, 6.36 m | Correct — 4.14 kW, 6.36 m |
| 6 | Impossible spec | Correct — refused | Correct — refused |
Six for six, both. Where the numbers differ at all it is in the third significant figure, traceable to which specific heat each one picked (4.184 versus 4.195 kJ/kg·K). Every trap was caught by both models.
So the test failed at its stated purpose. I could not break either one. What it surfaced instead is more interesting, and none of it is visible in the score column.
The real difference: what they say when you haven't asked
The 11% error nobody asked about
Problem 3 supplies a heating value of 45 MJ/Nm³. Both models used it and both got the arithmetic right. Only one questioned the number itself.
Delivered 407 kW / 32.6 Nm³/h / 31.7 m³/h, then flagged two things I had not asked about.
HHV vs LHV. Japanese 13A city gas is normally quoted at 45 MJ/Nm³ higher heating value, with LHV around 40.6. If the burner is rated on LHV input, true consumption is 36.1 Nm³/h — about 11% higher.
Rating basis. If 350,000 kcal/h is the output rather than the input, divide by burner efficiency and demand rises to roughly 38 Nm³/h.
It also noted that the near-cancellation of the two corrections here (×1.165 for temperature, ×0.835 for pressure) is "a coincidence worth being suspicious of, not a rule" — at 50 kPaG the same gas reads 27.5 m³/h, a 16% gap.
Delivered 406.8 kW / 32.54 Nm³/h / 31.66 m³/h with a clean derivation: normal conditions stated explicitly (0 °C, 101.325 kPa), combined gas law applied correctly, every step shown.
It flagged one ambiguity — the thermochemical calorie (4.184 kJ) versus the International Steam Table calorie (4.1868 kJ), a 0.07% difference.
It did not question whether 45 MJ/Nm³ was HHV or LHV.

Both answers are correct as asked. But an 11% error in gas demand sizes the wrong regulator and the wrong line. The calorie definition moves the answer by 0.07%. One of those caveats will cost you a site visit.
The area you would actually buy
Problem 4 asked for the required area, and 13.3 m² is the right answer to the question. It is not the number that goes on a purchase order.
After giving 13.3 m², it kept going.
- A real 1-shell-pass / 2-tube-pass unit is not pure counter-current. With P = 0.372 and R = 1.49, F ≈ 0.86 → 15.4 m²
- With typical fouling factors on both sides, service U drops to about 800 W/m²K → ≈ 21 m². In its words: "that's the number I'd actually buy — the clean-condition 13.3 m² is what the exchanger achieves on day one, not what it needs on year three"
- Cooling water returning at 55.4 °C is high for an open tower loop. Above roughly 50 °C carbonate scaling accelerates, which is exactly what drives that fouling factor up
Gave 13.25 m² with a complete and correct derivation: duty → cooling water outlet → LMTD → area, each step shown with units and the counter-current assumption stated.
It stopped at the question as asked.
The gap between 13.3 m² and 21 m² is the gap between a textbook answer and a specification.
Refuse and ask, or refuse and re-scope
Problem 6 was the one I expected to separate them, and it did — though not along the axis of right and wrong. Both models identified the thermodynamic impossibility immediately, and neither produced a fake design.
Opened with the blocker — 28 °C is unreachable with 32 °C cooling water, and the practical floor is 38–40 °C on a 6–8 K approach — then re-scoped and finished the job.
- Sized the unit to 40 °C: 719 kW total, about 62 m³/h of cooling water at a 10 K rise, roughly 10 m² split between condensing and subcooling zones
- Observed that subcooling is only 19% of the duty but eats two-thirds of the area, and suggested splitting the shell rather than oversizing one
- Then priced the original requirement anyway: if 28 °C is genuinely hard, a chilled-water trim cooler is 13.9 kW and under 0.5 m² — "the real cost is the 14 kW chiller"
- Finally asked why 28 °C, noting that if the reason is preventing flash in the condensate return, 40 °C already does it and the chiller is wasted capex
Refused cleanly and explained why: "It is physically impossible to cool the condensate below the inlet temperature of your cooling medium," adding that a realistic floor is 35–37 °C on a 3–5 K approach.
It then calculated the process-side duty (733 kW), since that depends only on the steam properties, and stopped to ask: revise the target to 40 °C, or supply a colder medium such as chilled water?
One small inconsistency: having called 28 °C impossible, it computed that duty to a 28 °C endpoint.

Which behaviour is better genuinely depends on you. Gemini's is the safer default — it does not assume what you meant. Claude's saves a round trip, and in this case it also answered the question I should have asked, which was whether 28 °C had ever been a real requirement.
Both had the context. One used it.
I ran all six problems as a single continuous conversation with each model — the way you actually work, not the way you run a clean benchmark. That means both models reached problem 6 carrying the same five problems' worth of context.
Only one of them reached back for it:
At a 10 K rise (32 → 42 °C, safely under the scaling threshold we discussed)
Given the earlier burner question, a combustion-air or feedwater preheat tie-in might be sitting right there
That second line is a heat-integration suggestion connecting two problems I had presented as unrelated. There is 140 kW of subcooling duty in problem 6 being thrown into cooling water, and a burner in problem 3 that needs combustion air. Tying them together is what a process engineer does after walking the whole plant — and I had not asked for it, in either problem.
Gemini answered each question completely and correctly, and treated each as its own problem. Same thread, same available context, different instinct about what to do with it.
Where Claude was sloppy
In problem 4, Claude wrote that "the streams cross" while describing the temperature profile. They do not. Cooling water leaves at 55.4 °C and the hot stream leaves at 60 °C — a close approach, not a cross. The F correction it applied was right and the area was right. The sentence was not.
It is a small thing, but it is the kind of small thing that matters. If you are skimming for confirmation rather than reading, that phrasing sends you looking for a problem that isn't there.
What this means if you do this work
Both models can do the arithmetic. At this level of difficulty that question is settled. Six problems, four deliberate traps, and not one wrong number between them. If you are still hand-carrying every unit conversion because you don't trust these tools to convert kcal/h to kW, you can stop.
The differentiator is domain-adjacent judgement, not calculation. HHV versus LHV, fouling allowance, the F correction, why the customer asked for 28 °C — none of it was in the question, and all of it changes what you buy.
Neither has earned an unchecked number. Every figure that leaves my desk still gets a hand calculation, and nothing here changes that. What these models have earned is the twenty minutes in front of the calculation: pulling the right correlation, laying out the assumptions, catching that I forgot the elevation term.
How I actually use them
I use Claude for nearly all of it. Not because it is more accurate — this test says it isn't — but because it surfaces the points I will have to deal with anyway, which means fewer trips back. When a model tells me up front that my heating value might be HHV, I resolve it then, rather than after the line sizing is finished.
That is a preference formed over months of real work, and you should read this article knowing it. If your workflow rewards a model that answers exactly what was asked and then stops, Gemini 3.1 Pro did that flawlessly here, and the tighter output is easier to paste into a calculation sheet.
Method and limitations
- The settings were not perfectly matched. Claude ran with thinking effort set to high; Gemini ran at its app default. Both scored 6/6, so this did not change the accuracy result — but the volume of unrequested caveats, which is where the real difference showed up, may partly reflect that setting. A matched-effort rerun is the obvious next test, and I intend to run it.
- One continuous conversation per model, not six isolated trials. Both models had the same accumulated context, so the comparison is like-for-like on that axis — but later answers are not independent of earlier ones, and Claude's demonstrably were not. A clean per-problem comparison needs fresh sessions for every question.
- Six problems, one run each, no retries. These models are non-deterministic and your results will differ. This is one engineer's field report, not a benchmark.
- A seventh problem — reading a scanned data sheet — was dropped. I could not use a real project document, and building a synthetic one would have tested something other than what I wanted to test.
- Both models are updated continuously. Results are specific to the versions and the date in the setup table.
- I have no commercial relationship with either vendor, and there are no affiliate links in this article.
FAQ
Do they handle unit conversion reliably? In this test, yes — including the case built to catch it. Problem 3 required converting kcal/h to kW and normal cubic metres to actual cubic metres at 45 °C and 20 kPaG, where the temperature correction (×1.165) and the pressure correction (×0.835) nearly cancel. Both models returned 31.7 m³/h and both got the direction right. Inverting that ratio is the classic error, and neither made it.
Can I use these for calculations that go into a stamped or certified deliverable? Not as the source. Any number carrying professional liability needs a hand calculation and a citable reference behind it. These models are good at getting you to the calculation and at catching what you forgot, and that is where the time actually goes.
Which one is better for long derivations? They were equally reliable here. Gemini's output is more consistently structured — numbered steps, every formula in display form — which is easier to transcribe into a calculation sheet. Claude's is more compressed and tends to lead with the answer, then discuss what could go wrong with it.