Key Takeaways
- →Claude Opus 4.8 scored 96.0% on 101 firearms questions, ahead of Claude Opus 5 at 93.1%, Claude Sonnet 5 at 91.1%, Gemini 3.1 Pro preview at 89.9%, Gemini 3.5 Flash at 88.9%, GPT-5.4 at 87.1%, Claude Haiku 4.5 at 78.2%, and GPT-4o-mini at 69.3%.
- →Post-2024 legal recency is the failure mode. Gemini 3.1 Pro preview answered 96.4% of everything else correctly and 53.3% of the recency questions it received. Five of eight models said the 2026 suppressor transfer tax is $200. It is $0.
- →Mainstream hardware is effectively solved. Six of eight models scored 100% on the 15 ballistics questions and 100% on every computation-tagged item; four scored 100% on the 13 AR-15 compatibility questions.
- →Nobody cleared 82% on boutique gear. The 11-question accessories set, seeded with products that do not exist, is the only category where no model reached 90%. Both Opus models topped it at 81.8%.
- →The newer flagship scored lower. Claude Opus 5 missed every question Claude Opus 4.8 missed, plus three more, on the same set under the same grader.
The Leaderboard
Claude Opus 4.8 leads the set at 96.0% full-credit accuracy, and the six models above 87% are separated by just under nine points. That compression is the first real finding: on questions about hardware, fitment, and ballistics, the current frontier is a crowded field rather than a leaderboard with a winner.
Full-credit accuracy counts only items scored 1.0, so a build plan that satisfies three of four criteria contributes nothing to this column. Every run below was graded under one identical rule set.
| Model | Accuracy | Ex-Recency | Items | Run |
|---|---|---|---|---|
| Claude Opus 4.8 | 96.0% | 96.5% | 101 | Jul 2, 2026 |
| Claude Opus 5 | 93.1% | 94.1% | 101 | Aug 17, 2026 |
| Claude Sonnet 5 | 91.1% | 91.8% | 101 | Aug 17, 2026 |
| Gemini 3.1 Pro preview | 89.9% | 96.4% | 99 | Jul 2, 2026 |
| Gemini 3.5 Flash | 88.9% | 94.1% | 99 | Jul 2, 2026 |
| GPT-5.4 | 87.1% | 89.4% | 101 | Jul 2, 2026 |
| Claude Haiku 4.5 | 78.2% | 81.2% | 101 | Aug 17, 2026 |
| GPT-4o-mini | 69.3% | 71.8% | 101 | Jul 2, 2026 |
The two Gemini runs answered 99 of the 101 items. Two questions errored during their July 2026 run and are excluded from both the numerator and the denominator of their scores.
The Ex-Recency column drops the post-2024 recency-tagged items, because a question about a 2026 tax change partly measures how recently a model was trained rather than how well it knows the domain. Read that column and the ordering below first place compresses: Gemini 3.1 Pro preview at 96.4% is a statistical tie with Claude Opus 4.8 at 96.5%. The full-set number is still the one that matters to a user asking a deployed model a question today, but the two columns measure different things and the gap between them is mostly training cutoff.
Category Scores Separate the Models, the Topline Does Not
Two categories account for nearly all the spread: legal and NFA questions, where the ceiling is 95.2% and the floor is 57.1%, and boutique accessories, where the ceiling is 81.8%. Everything else is close to saturated. Six of eight models score 100% on ballistics, and four score 100% on AR-15 compatibility.
| Category | Opus 4.8 | Opus 5 | Sonnet 5 | G3.1 Pro | G3.5 Flash | GPT-5.4 | Haiku 4.5 | 4o-mini |
|---|---|---|---|---|---|---|---|---|
| Optics footprints (12) | 100 | 91.7 | 91.7 | 100 | 100 | 91.7 | 75 | 66.7 |
| Suppressors (9) | 88.9 | 88.9 | 100 | 100 | 100 | 100 | 77.8 | 88.9 |
| AR-15 compatibility (13) | 100 | 92.3 | 100 | 100 | 92.3 | 100 | 84.6 | 76.9 |
| Handgun platforms (10) | 100 | 100 | 100 | 88.9 | 88.9 | 100 | 90 | 100 |
| Ballistics (15) | 100 | 100 | 100 | 100 | 100 | 100 | 86.7 | 66.7 |
| Legal / NFA (21) | 95.2 | 90.5 | 90.5 | 71.4 | 75 | 81 | 76.2 | 57.1 |
| Build planning (6) | 100 | 100 | 96.7 | 100 | 93.3 | 77.2 | 80 | 87.8 |
| Boutique accessories (11) | 81.8 | 81.8 | 72.7 | 72.7 | 72.7 | 63.6 | 72.7 | 54.5 |
| Competition rules (4) | 100 | 100 | 50 | 100 | 100 | 75 | 100 | 50 |
Values are score-weighted, so build plans contribute the fraction of criteria they satisfied. Item counts in parentheses are for the full 101-question set; each Gemini run is missing two items. Both lack one handgun platforms question, Gemini 3.1 Pro preview also lacks one AR-15 compatibility question, and Gemini 3.5 Flash one legal question.
Competition rules is the noisiest row because it holds only four questions, so a single miss moves it 25 points. Claude Sonnet 5 and GPT-4o-mini both landed at 50% there. Sonnet 5 put the USPSA Production magazine capacity limit at 10 rounds instead of 15, and placed a frame-mounted red dot on a Glock 34 in Limited Optics when a frame mount pushes the gun to Open. GPT-5.4 made the same 10-round call. Our USPSA Carry Optics guide covers where each division line actually falls.
Recency Is the Dangerous Failure, Not Trivia
Sixteen of the 101 questions turn on something that changed after 2024, and that is where the models come apart. Gemini 3.1 Pro preview answered 96.4% of every other question correctly and 53.3% of the recency set. Gemini 3.5 Flash ran 94.1% against 57.1%. GPT-5.4 ran 89.4% against 75.0%. Claude Opus 4.8 has the smallest gap in the study, 96.5% against 93.8%. The items that errored out of the Gemini runs include recency questions, so their recency denominators are 15 and 14 rather than 16.

The specific misses matter more than the percentages, because these are the questions a real buyer asks. Asked what the 2026 federal transfer tax is to acquire a suppressor on a Form 4, five of the eight models answered $200. The correct answer is $0: the One Big Beautiful Bill Act zeroed the making and transfer tax on suppressors, short-barreled rifles, short-barreled shotguns, and any other weapons effective 2026, while leaving the Form 4 itself, fingerprints, and NFA registration fully in place. Our suppressor buying guide walks the current process. Five models made the same $200 call on the Form 1 short-barreled rifle making tax, and four said the AOW transfer tax is still the historical $5 stamp.
The court question produced a cleaner failure. Asked what the Fifth Circuit held about suppressors in its 2026 decision, Claude Opus 5, Gemini 3.1 Pro preview, and Gemini 3.5 Flash all selected the exact inversion: that suppressors are not arms and receive no constitutional protection. The panel held the opposite in United States v. Comeaux, finding on June 18, 2026 that suppressors are Arms within the plain text of the Second Amendment while still affirming the underlying NFA conviction.
Forced reset triggers produced a third. Two models put the 2025 DOJ and Rare Breed settlement's carve-out on the FN PS90 rather than on grip-fed striker pistols such as the Glock, M&P, and Canik, where the magazine loads through the trigger hand. That distinction decides whether a given trigger is settled federally or still awaiting classification, and it sits on top of a separate layer of state law that operates independently of the federal position.
None of these answers arrived hedged. A model that is stale on a tax rate states the stale rate in a complete, confident sentence, which is precisely what makes recency failure worse than ordinary ignorance. An assistant that says it does not know sends the reader to ATF or a dealer. An assistant that says $200 sends the reader away satisfied and wrong.
The Question Six of Eight Models Failed the Same Way
Asked for the legal barrel length after permanently pinning and welding a 1.5-inch muzzle device to a 14.5-inch barrel, six of the eight models answered 16.0 inches. Every one of the six picked the same wrong option. The correct answer is 15.5 inches, because a muzzle device that threads onto the barrel does not add its full overall length: the threaded section overlaps the muzzle, so the net gain is shorter than the device.
This is the most consequential miss in the study. ATF measures barrel length as the barrel plus any permanently attached muzzle device, with a blind pin at least 0.125 inch deep. A builder who follows the 16.0-inch answer assembles what they believe is a legal rifle and actually holds an unregistered short-barreled rifle. The two Gemini models were the only ones to get it right.
The related arithmetic question, where a 2.0-inch device threads onto 0.5 inch of a 14.5-inch barrel for a measured 16.0 inches, was answered correctly by seven of eight. The models can do the subtraction when the overlap is stated. They fail when the overlap has to be inferred, which is exactly how the question arrives in real life.
What Happens When the Product Does Not Exist
Several questions in the set describe products that do not exist, and score the model on refusing the bait rather than on recall. The boutique accessories category is built around them, and it is the only category in the study where no model reached 90%. Both Opus models topped it at 81.8%, GPT-5.4 landed at 63.6%, and GPT-4o-mini at 54.5%.
The Henning Group question is the clearest case. Henning makes magazine basepads and small competition parts, not magwells, so the correct answer to "which magwell does Henning make for the CZ Shadow 2" is that no such product exists. GPT-5.4 and GPT-4o-mini both selected an invented Henning GrandMaster flared magwell. On the open-ended version of the same question, Gemini 3.5 Flash and GPT-4o-mini both described Henning as manufacturing magwells alongside basepads. Every Claude model declined the bait on both versions.
The only question in the set that every model missed is the reverse: which platform has no production forced reset trigger available. The answer is the CZ Scorpion EVO, where the announced product has never shipped. All eight models got it wrong, and five of them named the FN PS90 instead. A model asked to identify absence reaches for whichever option feels most obscure rather than checking what is actually purchasable.
Material specifications on small-shop parts failed the same way. Six of eight models, including the top four scorers, missed the trip material on one small-batch Canik forced reset trigger, splitting between 17-4 PH stainless and Grade 5 titanium when the maker publishes 316 stainless. That is the shape of the failure across this whole category: the model produces a material that sounds correct for the part class instead of the one the manufacturer lists.
Pistol Red Dots in the Footprint Questions
Affiliate links (?)
Footprint Questions Are the Quiet Weak Spot
Optics footprints look solved at the top and are not. Three models scored 100% on the 12-question set, and three more sat one question back at 91.7%, but the misses cluster on the questions a buyer actually has. Claude Opus 5 answered that a Glock 43X MOS natively fits a Holosun EPS Carry, when the 43X MOS is cut for the standard Shield RMSc footprint and the EPS Carry uses the modified Holosun K pattern. That is an optic ordered for a slide it does not bolt to.
Mount height produced a similar split. Asked for the standard absolute co-witness height for an AR-15 red dot, measured from the top of the receiver rail to the optical axis, Claude Sonnet 5 and GPT-5.4 both answered 1.57 inches and Claude Haiku 4.5 answered 1.93 inches. The correct value is 1.41 inches. The two wrong answers are real mount heights: 1.54 to 1.70 inches is the lower-third band, and 1.93 inches is a heads-up height where standard irons drop out of the window entirely.
Both Gemini models missed the Glock Gen6 optic-mounting question, answering that Gen6 carried the MOS system forward from Gen5 rather than introducing a new cut that ships with a factory polymer plate. The distinction is the whole reason a Gen5 plate does not fit a Gen6, which our Glock Gen6 optic plate guide works through. Suppressor mounts drew four misses on a single question, with Claude Opus 4.8, Claude Opus 5, Claude Haiku 4.5, and GPT-4o-mini all attributing the KeyMo quick-detach system to SureFire or SilencerCo rather than Dead Air.
The Newest Claude Flagship Scored Below Its Predecessor
Claude Opus 4.8 scored 96.0% and Claude Opus 5 scored 93.1% on the same 101 questions under the same grader. Opus 5 missed every item Opus 4.8 missed, then added three: the Glock 43X MOS and EPS Carry footprint question, AR-15 gas ring orientation, and the Fifth Circuit suppressor holding. There is no item Opus 4.8 got wrong and Opus 5 got right.
Three questions on a 101-item set is a narrow margin, and this study is not powered to call a general capability regression from it. The two runs also sat six weeks apart, on July 2 and August 17, 2026. What the result does establish is that a newer flagship is not automatically better on niche domain recall, and that a domain-specific benchmark is the only way to find out which one your use case wants.
The four questions both Opus models missed are worth naming, because they mark the current edge of the domain: the pin-and-weld barrel length, the KeyMo mount attribution, the small-shop trigger material, and the CZ Scorpion EVO forced reset trigger. The Scorpion question defeated all eight models, and the pin-and-weld and trigger-material questions each defeated six of eight.
Methodology and Disclosure
The benchmark is 101 questions: 89 multiple choice, 6 free text, and 6 build plans, spread across optics footprints, AR-15 compatibility, handgun platforms, suppressors, ballistics, legal and NFA, USPSA competition, boutique accessories, and build planning. Eighty-eight of the 101 are tagged hard, and 16 turn on something that changed after 2024. Every answer key traces to our published guides and articles, our product catalog, or our internal firearm fact ledger.
Multiple-choice options are reshuffled per question before the model sees them, using a deterministic shuffle seeded by question ID, so no model can exploit the position an answer was authored in. The default seed produces a presented-key distribution of 22 A, 22 B, 22 C, and 23 D across the 89 multiple-choice items. Free-text and build-plan answers are graded by a cross-family LLM judge, GPT-5.4-mini, chosen so that no model family grades its own answers. Deterministic substring matching runs first as an accept-only fast path, and build plans are graded criterion by criterion, with each criterion judged in isolation so an extra-rubric detail cannot fail a build.
The grader changed between the first runs and this publication. The original version auto-failed any free-text answer containing a trap term, which penalized answers that correctly negated it, as in "$0, the old $200 tax was zeroed" or "Henning does not make magwells." That is exactly the anti-hallucination behavior the set is meant to reward. Under the current grader the match block is accept-only and anything that does not pass cleanly goes to the strict judge instead of failing automatically. Every published run in this article was re-graded from its stored raw responses under that single rule set.
The models were not all reached the same way. GPT-5.4, GPT-4o-mini, Gemini 3.1 Pro preview, and Gemini 3.5 Flash were called question by question through their APIs at temperature 0. The four Claude models were run as tool-forbidden Claude Code subagents answering in four batches of roughly 25 questions each, with harness-default decoding rather than an explicit temperature-0 setting, because that was the access we had. Both paths received the same shuffled options and the same grading. Batched answering gives the Claude runs the preceding questions in context, which the per-question API calls did not have, and the harness defaults are a second difference. Treat cross-family gaps of a few points as provisional for that reason.
One further disclosure. This benchmark was orchestrated by Claude-based agents, this article was drafted by one, and four of the eight models under test are Anthropic models. That is why the grading path is deterministic matching plus a cross-family judge rather than a same-family evaluator, why every question ID, expected answer, and given answer is recorded in the run logs, and why the result that reflects worst on the newest Anthropic flagship is reported in its own section rather than buried.
Three limits travel with the numbers. The recency items partly measure training-data cutoff rather than domain knowledge, which is why the leaderboard reports an ex-recency column, and every model was tested on parametric knowledge alone: deployed chat products with web search enabled may answer current-law questions correctly even where their base model tested stale here. A future edition will also score with negative marking and an explicit abstention option, so a model that says to verify current law outscores one that states a stale figure confidently. The Gemini runs answered 99 of 101 items because two questions errored during their July 2026 run, so their scores use a 99-item denominator. And this is a point-in-time snapshot: the runs are dated July 2 and August 17, 2026, model versions are replaced on a scale of months, and firearm law moves at least as fast. A model that fails a recency question today can pass it after its next training refresh, and a legal answer that is correct today can be wrong by the next court term. More published data sets live at Rifle Configurator Research.
How to Use an AI Assistant for Gun Questions
Use it for hardware, verify anything legal. The data supports a clean split: mainstream ballistics, gas systems, buffer weights, cartridge lineage, and mount arithmetic are close to saturated across every frontier model, and those answers can be trusted at roughly the rate a knowledgeable forum answer can. Tax rates, court holdings, settlement scope, state statutes, and whether a given part exists are not.
The practical rule is that the confidence of the answer carries no information about its accuracy. A model stating that the suppressor stamp is $200, that the Fifth Circuit denied constitutional protection to suppressors, or that a boutique maker sells a magwell it has never produced does so in the same tone it uses for a twist rate it has right. For fitment specifically, checking a claim against a compatibility source costs less than a returned optic; our rifle builder enforces the same footprint, thread, and generation gates the models missed.
Get the Next Benchmark
We publish original firearms research, builder-behavior data, and hands-on reviews. Subscribe for the next data set.
















