We're on Instagram. Follow @rifleconfigurator
Home/Articles/Industry
Industry

Which AI Model Knows Guns Best? 101-Question Benchmark

We scored eight AI models on 101 firearms questions covering optics footprints, AR-15 compatibility, suppressors, ballistics, NFA law, USPSA rules, and boutique gear. Claude Opus 4.8 led at 96.0%, but every model failed on post-2024 legal recency, and five of eight still put the 2026 suppressor transfer tax at $200.

Author
AB
Read
13 min
Which AI Model Knows Guns Best? 101-Question Benchmark header image

Key Takeaways

  • Claude Opus 4.8 scored 96.0% on 101 firearms questions, ahead of Claude Opus 5 at 93.1%, Claude Sonnet 5 at 91.1%, Gemini 3.1 Pro preview at 89.9%, Gemini 3.5 Flash at 88.9%, GPT-5.4 at 87.1%, Claude Haiku 4.5 at 78.2%, and GPT-4o-mini at 69.3%.
  • Post-2024 legal recency is the failure mode. Gemini 3.1 Pro preview answered 96.4% of everything else correctly and 53.3% of the recency questions it received. Five of eight models said the 2026 suppressor transfer tax is $200. It is $0.
  • Mainstream hardware is effectively solved. Six of eight models scored 100% on the 15 ballistics questions and 100% on every computation-tagged item; four scored 100% on the 13 AR-15 compatibility questions.
  • Nobody cleared 82% on boutique gear. The 11-question accessories set, seeded with products that do not exist, is the only category where no model reached 90%. Both Opus models topped it at 81.8%.
  • The newer flagship scored lower. Claude Opus 5 missed every question Claude Opus 4.8 missed, plus three more, on the same set under the same grader.

The Leaderboard

Claude Opus 4.8 leads the set at 96.0% full-credit accuracy, and the six models above 87% are separated by just under nine points. That compression is the first real finding: on questions about hardware, fitment, and ballistics, the current frontier is a crowded field rather than a leaderboard with a winner.

Full-credit accuracy counts only items scored 1.0, so a build plan that satisfies three of four criteria contributes nothing to this column. Every run below was graded under one identical rule set.

Full-credit accuracy by model on the 101-question firearms knowledge benchmark, July and August 2026.
ModelAccuracyEx-RecencyItemsRun
Claude Opus 4.896.0%96.5%101Jul 2, 2026
Claude Opus 593.1%94.1%101Aug 17, 2026
Claude Sonnet 591.1%91.8%101Aug 17, 2026
Gemini 3.1 Pro preview89.9%96.4%99Jul 2, 2026
Gemini 3.5 Flash88.9%94.1%99Jul 2, 2026
GPT-5.487.1%89.4%101Jul 2, 2026
Claude Haiku 4.578.2%81.2%101Aug 17, 2026
GPT-4o-mini69.3%71.8%101Jul 2, 2026

The two Gemini runs answered 99 of the 101 items. Two questions errored during their July 2026 run and are excluded from both the numerator and the denominator of their scores.

The Ex-Recency column drops the post-2024 recency-tagged items, because a question about a 2026 tax change partly measures how recently a model was trained rather than how well it knows the domain. Read that column and the ordering below first place compresses: Gemini 3.1 Pro preview at 96.4% is a statistical tie with Claude Opus 4.8 at 96.5%. The full-set number is still the one that matters to a user asking a deployed model a question today, but the two columns measure different things and the gap between them is mostly training cutoff.

Category Scores Separate the Models, the Topline Does Not

Two categories account for nearly all the spread: legal and NFA questions, where the ceiling is 95.2% and the floor is 57.1%, and boutique accessories, where the ceiling is 81.8%. Everything else is close to saturated. Six of eight models score 100% on ballistics, and four score 100% on AR-15 compatibility.

Score-weighted accuracy by question category and model. Build plans contribute partial credit.
CategoryOpus 4.8Opus 5Sonnet 5G3.1 ProG3.5 FlashGPT-5.4Haiku 4.54o-mini
Optics footprints (12)10091.791.710010091.77566.7
Suppressors (9)88.988.910010010010077.888.9
AR-15 compatibility (13)10092.310010092.310084.676.9
Handgun platforms (10)10010010088.988.910090100
Ballistics (15)10010010010010010086.766.7
Legal / NFA (21)95.290.590.571.4758176.257.1
Build planning (6)10010096.710093.377.28087.8
Boutique accessories (11)81.881.872.772.772.763.672.754.5
Competition rules (4)100100501001007510050

Values are score-weighted, so build plans contribute the fraction of criteria they satisfied. Item counts in parentheses are for the full 101-question set; each Gemini run is missing two items. Both lack one handgun platforms question, Gemini 3.1 Pro preview also lacks one AR-15 compatibility question, and Gemini 3.5 Flash one legal question.

Competition rules is the noisiest row because it holds only four questions, so a single miss moves it 25 points. Claude Sonnet 5 and GPT-4o-mini both landed at 50% there. Sonnet 5 put the USPSA Production magazine capacity limit at 10 rounds instead of 15, and placed a frame-mounted red dot on a Glock 34 in Limited Optics when a frame mount pushes the gun to Open. GPT-5.4 made the same 10-round call. Our USPSA Carry Optics guide covers where each division line actually falls.

Recency Is the Dangerous Failure, Not Trivia

Sixteen of the 101 questions turn on something that changed after 2024, and that is where the models come apart. Gemini 3.1 Pro preview answered 96.4% of every other question correctly and 53.3% of the recency set. Gemini 3.5 Flash ran 94.1% against 57.1%. GPT-5.4 ran 89.4% against 75.0%. Claude Opus 4.8 has the smallest gap in the study, 96.5% against 93.8%. The items that errored out of the Gemini runs include recency questions, so their recency denominators are 15 and 14 rather than 16.

Grouped bar chart comparing each model's full-credit accuracy on post-2024 recency questions against all other question types. Gemini 3.1 Pro preview shows the widest gap, 96.4 percent on other questions against 53.3 percent on recency.
Every model scores worse on post-2024 recency than on everything else. The gap ranges from 2.7 points for Claude Opus 4.8 to 43.1 points for Gemini 3.1 Pro preview (Source: Rifle Configurator Research)

The specific misses matter more than the percentages, because these are the questions a real buyer asks. Asked what the 2026 federal transfer tax is to acquire a suppressor on a Form 4, five of the eight models answered $200. The correct answer is $0: the One Big Beautiful Bill Act zeroed the making and transfer tax on suppressors, short-barreled rifles, short-barreled shotguns, and any other weapons effective 2026, while leaving the Form 4 itself, fingerprints, and NFA registration fully in place. Our suppressor buying guide walks the current process. Five models made the same $200 call on the Form 1 short-barreled rifle making tax, and four said the AOW transfer tax is still the historical $5 stamp.

The court question produced a cleaner failure. Asked what the Fifth Circuit held about suppressors in its 2026 decision, Claude Opus 5, Gemini 3.1 Pro preview, and Gemini 3.5 Flash all selected the exact inversion: that suppressors are not arms and receive no constitutional protection. The panel held the opposite in United States v. Comeaux, finding on June 18, 2026 that suppressors are Arms within the plain text of the Second Amendment while still affirming the underlying NFA conviction.

Forced reset triggers produced a third. Two models put the 2025 DOJ and Rare Breed settlement's carve-out on the FN PS90 rather than on grip-fed striker pistols such as the Glock, M&P, and Canik, where the magazine loads through the trigger hand. That distinction decides whether a given trigger is settled federally or still awaiting classification, and it sits on top of a separate layer of state law that operates independently of the federal position.

None of these answers arrived hedged. A model that is stale on a tax rate states the stale rate in a complete, confident sentence, which is precisely what makes recency failure worse than ordinary ignorance. An assistant that says it does not know sends the reader to ATF or a dealer. An assistant that says $200 sends the reader away satisfied and wrong.

The Question Six of Eight Models Failed the Same Way

Asked for the legal barrel length after permanently pinning and welding a 1.5-inch muzzle device to a 14.5-inch barrel, six of the eight models answered 16.0 inches. Every one of the six picked the same wrong option. The correct answer is 15.5 inches, because a muzzle device that threads onto the barrel does not add its full overall length: the threaded section overlaps the muzzle, so the net gain is shorter than the device.

This is the most consequential miss in the study. ATF measures barrel length as the barrel plus any permanently attached muzzle device, with a blind pin at least 0.125 inch deep. A builder who follows the 16.0-inch answer assembles what they believe is a legal rifle and actually holds an unregistered short-barreled rifle. The two Gemini models were the only ones to get it right.

The related arithmetic question, where a 2.0-inch device threads onto 0.5 inch of a 14.5-inch barrel for a measured 16.0 inches, was answered correctly by seven of eight. The models can do the subtraction when the overlap is stated. They fail when the overlap has to be inferred, which is exactly how the question arrives in real life.

What Happens When the Product Does Not Exist

Several questions in the set describe products that do not exist, and score the model on refusing the bait rather than on recall. The boutique accessories category is built around them, and it is the only category in the study where no model reached 90%. Both Opus models topped it at 81.8%, GPT-5.4 landed at 63.6%, and GPT-4o-mini at 54.5%.

The Henning Group question is the clearest case. Henning makes magazine basepads and small competition parts, not magwells, so the correct answer to "which magwell does Henning make for the CZ Shadow 2" is that no such product exists. GPT-5.4 and GPT-4o-mini both selected an invented Henning GrandMaster flared magwell. On the open-ended version of the same question, Gemini 3.5 Flash and GPT-4o-mini both described Henning as manufacturing magwells alongside basepads. Every Claude model declined the bait on both versions.

The only question in the set that every model missed is the reverse: which platform has no production forced reset trigger available. The answer is the CZ Scorpion EVO, where the announced product has never shipped. All eight models got it wrong, and five of them named the FN PS90 instead. A model asked to identify absence reaches for whichever option feels most obscure rather than checking what is actually purchasable.

Material specifications on small-shop parts failed the same way. Six of eight models, including the top four scorers, missed the trip material on one small-batch Canik forced reset trigger, splitting between 17-4 PH stainless and Grade 5 titanium when the maker publishes 316 stainless. That is the shape of the failure across this whole category: the model produces a material that sounds correct for the part class instead of the one the manufacturer lists.

Pistol Red Dots in the Footprint Questions

Osight XE AMRS Enclosed (RMR Footprint) product image
Pistol Optics • $249.99

Osight XE AMRS Enclosed (RMR Footprint)

  • AMRS: 5 reticles (2/6 MOA dot, 32 MOA circle)
  • Enclosed emitter, RMR footprint
$249.99 Catalog
View at Amazon
Osight XR Enclosed (RMR Footprint) product image
Pistol Optics • $299.99

Osight XR Enclosed (RMR Footprint)

  • 2 MOA / 6 MOA / 32 MOA MRS
  • Enclosed emitter, RMR footprint
$299.99 Catalog
View at Amazon
Osight SE Enclosed (6 MOA Red) product image
Pistol Optics • $179.99

Osight SE Enclosed (6 MOA Red)

  • 6 MOA red dot
  • Enclosed emitter, RMSc footprint
$179.99 Catalog
View at Amazon
Osight SE DPP Enclosed Red Dot product image
Pistol Optics • $239.99

Osight SE DPP Enclosed Red Dot

  • 2 MOA dot + 32 MOA circle MRS
  • Enclosed emitter, DeltaPoint Pro footprint
$239.99 Catalog
View at OpticsPlanet
Osight SE Enclosed (2 MOA + 32 MOA Red MRS) product image
Pistol Optics • $199.99

Osight SE Enclosed (2 MOA + 32 MOA Red MRS)

  • 2 MOA dot + 32 MOA circle MRS
  • Enclosed emitter, RMSc footprint
$199.99 Catalog
View at Amazon
SIG ROMEO-X Compact Enclosed product image
Pistol Optics • $499.99

SIG ROMEO-X Compact Enclosed

  • 3 MOA / 6 MOA / Circle Dot
  • Enclosed emitter
$499.99
View at OpticsPlanet

Affiliate links (?)

Scroll

Footprint Questions Are the Quiet Weak Spot

Optics footprints look solved at the top and are not. Three models scored 100% on the 12-question set, and three more sat one question back at 91.7%, but the misses cluster on the questions a buyer actually has. Claude Opus 5 answered that a Glock 43X MOS natively fits a Holosun EPS Carry, when the 43X MOS is cut for the standard Shield RMSc footprint and the EPS Carry uses the modified Holosun K pattern. That is an optic ordered for a slide it does not bolt to.

Mount height produced a similar split. Asked for the standard absolute co-witness height for an AR-15 red dot, measured from the top of the receiver rail to the optical axis, Claude Sonnet 5 and GPT-5.4 both answered 1.57 inches and Claude Haiku 4.5 answered 1.93 inches. The correct value is 1.41 inches. The two wrong answers are real mount heights: 1.54 to 1.70 inches is the lower-third band, and 1.93 inches is a heads-up height where standard irons drop out of the window entirely.

Both Gemini models missed the Glock Gen6 optic-mounting question, answering that Gen6 carried the MOS system forward from Gen5 rather than introducing a new cut that ships with a factory polymer plate. The distinction is the whole reason a Gen5 plate does not fit a Gen6, which our Glock Gen6 optic plate guide works through. Suppressor mounts drew four misses on a single question, with Claude Opus 4.8, Claude Opus 5, Claude Haiku 4.5, and GPT-4o-mini all attributing the KeyMo quick-detach system to SureFire or SilencerCo rather than Dead Air.

The Newest Claude Flagship Scored Below Its Predecessor

Claude Opus 4.8 scored 96.0% and Claude Opus 5 scored 93.1% on the same 101 questions under the same grader. Opus 5 missed every item Opus 4.8 missed, then added three: the Glock 43X MOS and EPS Carry footprint question, AR-15 gas ring orientation, and the Fifth Circuit suppressor holding. There is no item Opus 4.8 got wrong and Opus 5 got right.

Three questions on a 101-item set is a narrow margin, and this study is not powered to call a general capability regression from it. The two runs also sat six weeks apart, on July 2 and August 17, 2026. What the result does establish is that a newer flagship is not automatically better on niche domain recall, and that a domain-specific benchmark is the only way to find out which one your use case wants.

The four questions both Opus models missed are worth naming, because they mark the current edge of the domain: the pin-and-weld barrel length, the KeyMo mount attribution, the small-shop trigger material, and the CZ Scorpion EVO forced reset trigger. The Scorpion question defeated all eight models, and the pin-and-weld and trigger-material questions each defeated six of eight.

Methodology and Disclosure

The benchmark is 101 questions: 89 multiple choice, 6 free text, and 6 build plans, spread across optics footprints, AR-15 compatibility, handgun platforms, suppressors, ballistics, legal and NFA, USPSA competition, boutique accessories, and build planning. Eighty-eight of the 101 are tagged hard, and 16 turn on something that changed after 2024. Every answer key traces to our published guides and articles, our product catalog, or our internal firearm fact ledger.

Multiple-choice options are reshuffled per question before the model sees them, using a deterministic shuffle seeded by question ID, so no model can exploit the position an answer was authored in. The default seed produces a presented-key distribution of 22 A, 22 B, 22 C, and 23 D across the 89 multiple-choice items. Free-text and build-plan answers are graded by a cross-family LLM judge, GPT-5.4-mini, chosen so that no model family grades its own answers. Deterministic substring matching runs first as an accept-only fast path, and build plans are graded criterion by criterion, with each criterion judged in isolation so an extra-rubric detail cannot fail a build.

The grader changed between the first runs and this publication. The original version auto-failed any free-text answer containing a trap term, which penalized answers that correctly negated it, as in "$0, the old $200 tax was zeroed" or "Henning does not make magwells." That is exactly the anti-hallucination behavior the set is meant to reward. Under the current grader the match block is accept-only and anything that does not pass cleanly goes to the strict judge instead of failing automatically. Every published run in this article was re-graded from its stored raw responses under that single rule set.

The models were not all reached the same way. GPT-5.4, GPT-4o-mini, Gemini 3.1 Pro preview, and Gemini 3.5 Flash were called question by question through their APIs at temperature 0. The four Claude models were run as tool-forbidden Claude Code subagents answering in four batches of roughly 25 questions each, with harness-default decoding rather than an explicit temperature-0 setting, because that was the access we had. Both paths received the same shuffled options and the same grading. Batched answering gives the Claude runs the preceding questions in context, which the per-question API calls did not have, and the harness defaults are a second difference. Treat cross-family gaps of a few points as provisional for that reason.

One further disclosure. This benchmark was orchestrated by Claude-based agents, this article was drafted by one, and four of the eight models under test are Anthropic models. That is why the grading path is deterministic matching plus a cross-family judge rather than a same-family evaluator, why every question ID, expected answer, and given answer is recorded in the run logs, and why the result that reflects worst on the newest Anthropic flagship is reported in its own section rather than buried.

Three limits travel with the numbers. The recency items partly measure training-data cutoff rather than domain knowledge, which is why the leaderboard reports an ex-recency column, and every model was tested on parametric knowledge alone: deployed chat products with web search enabled may answer current-law questions correctly even where their base model tested stale here. A future edition will also score with negative marking and an explicit abstention option, so a model that says to verify current law outscores one that states a stale figure confidently. The Gemini runs answered 99 of 101 items because two questions errored during their July 2026 run, so their scores use a 99-item denominator. And this is a point-in-time snapshot: the runs are dated July 2 and August 17, 2026, model versions are replaced on a scale of months, and firearm law moves at least as fast. A model that fails a recency question today can pass it after its next training refresh, and a legal answer that is correct today can be wrong by the next court term. More published data sets live at Rifle Configurator Research.

How to Use an AI Assistant for Gun Questions

Use it for hardware, verify anything legal. The data supports a clean split: mainstream ballistics, gas systems, buffer weights, cartridge lineage, and mount arithmetic are close to saturated across every frontier model, and those answers can be trusted at roughly the rate a knowledgeable forum answer can. Tax rates, court holdings, settlement scope, state statutes, and whether a given part exists are not.

The practical rule is that the confidence of the answer carries no information about its accuracy. A model stating that the suppressor stamp is $200, that the Fifth Circuit denied constitutional protection to suppressors, or that a boutique maker sells a magwell it has never produced does so in the same tone it uses for a twist rate it has right. For fitment specifically, checking a claim against a compatibility source costs less than a returned optic; our rifle builder enforces the same footprint, thread, and generation gates the models missed.

Get the Next Benchmark

We publish original firearms research, builder-behavior data, and hands-on reviews. Subscribe for the next data set.

Free targets, drill cards, and weekly reviews by email. Follow us on Instagram and Facebook for daily builds and gear picks.

Follow

Frequently Asked Questions

Which AI model knows the most about guns?
Claude Opus 4.8 scored highest on our benchmark, answering 96.0 percent of 101 firearms questions correctly for full credit. Claude Opus 5 followed at 93.1 percent, Claude Sonnet 5 at 91.1 percent, Gemini 3.1 Pro preview at 89.9 percent, Gemini 3.5 Flash at 88.9 percent, GPT-5.4 at 87.1 percent, Claude Haiku 4.5 at 78.2 percent, and GPT-4o-mini at 69.3 percent. The top six are separated by just under nine points, so the practical answer is that any current frontier model handles mainstream firearms hardware well and none of them is reliable on 2026 legal specifics.
Can you trust ChatGPT, Claude, or Gemini for NFA and gun law questions?
No, not without checking the answer against a primary source. Legal and NFA questions were the second-weakest category in our benchmark, with a ceiling of 95.2 percent and a floor of 57.1 percent. Five of the eight models answered that the 2026 federal transfer tax on a suppressor is $200 when the correct answer is $0, because the One Big Beautiful Bill Act zeroed that tax effective 2026. Three models, including Claude Opus 5, inverted the Fifth Circuit's June 2026 holding in United States v. Comeaux and said the court found suppressors are not protected arms. A model that is stale on tax law states the stale figure with full confidence and no hedge.
How were the AI models tested?
Each model answered the same 101 questions: 89 multiple choice, 6 free text, and 6 build plans, covering optics footprints, AR-15 compatibility, handgun platforms, suppressors, ballistics, NFA and state law, USPSA competition rules, boutique aftermarket gear, and build planning. Multiple-choice options were reshuffled by a deterministic seeded shuffle so no model could exploit answer position, producing a presented-key distribution of 22 A, 22 B, 22 C, and 23 D across the 89 multiple-choice items. Free-text and build-plan answers were graded by a cross-family judge, GPT-5.4-mini, with deterministic substring matching as an accept-only fast path. Answer keys come from our published guides and articles, our product catalog, and our internal firearm fact ledger.
What is hallucination bait in an AI benchmark?
A hallucination-bait question offers a plausible-sounding product that does not exist and rewards the model for saying so. Our set includes several. One asks which magwell the Henning Group makes for the CZ Shadow 2; Henning makes magazine basepads and small competition parts, not magwells, and the correct answer is that no such product exists. GPT-5.4 and GPT-4o-mini both picked an invented Henning GrandMaster flared magwell. Another asks which platform has no production forced reset trigger available, where the answer is the CZ Scorpion EVO. All eight models got that one wrong, and five of them named the FN PS90 instead.
Did the newest Claude model score highest?
No. Claude Opus 4.8 scored 96.0 percent and Claude Opus 5 scored 93.1 percent on the same 101 questions under the same grading rules. Opus 5 missed every item Opus 4.8 missed plus three more: the Glock 43X MOS and Holosun EPS Carry footprint mismatch, AR-15 gas ring orientation, and the Fifth Circuit suppressor holding. Three questions on a 101-item set is a narrow margin and should not be read as a general capability regression, but it is what the runs show and we are not adjusting it.
Why do AI models get firearms law wrong so often?
Because firearms law moved faster than training data between 2024 and 2026 and models default to the figure that dominated their training corpus. The $200 NFA stamp appeared in decades of writing before the One Big Beautiful Bill Act zeroed the making and transfer tax on suppressors, short-barreled rifles, short-barreled shotguns, and any other weapons for 2026. On the post-2024 recency questions in its run, Gemini 3.1 Pro preview scored 53.3 percent while scoring 96.4 percent on everything else, the widest split in the study. The failure is not general weakness, it is a specific and dangerous staleness.
What did every model get right?
Ballistics and arithmetic. Six of the eight models scored 100 percent on the 15 ballistics questions, covering twist rates, cartridge lineage, MOA and mrad conversion, and barrel-length velocity, and only Claude Haiku 4.5 at 86.7 percent and GPT-4o-mini at 66.7 percent dropped points. The same six scored 100 percent on the computation-tagged items. Mainstream AR-15 compatibility is nearly as solved: four of the eight scored 100 percent on the 13-question set covering gas systems, buffer internals, and clone-correct specifications, and the lowest score in that category was 76.9 percent.
Share
Pass the dispatch