VigilSAR Defense LLM Benchmark — which models can be trusted with ISR work
AIThis post was created with the assistance of artificial intelligence (AI).
VigilSAR Defense LLM Benchmark
The public benchmark page — aggregate results public, task set private. Source: vigilsar.com

VIGILSAR, a company specializing in defense-ISR software, has taken a bold step in the artificial intelligence space by publishing a public leaderboard for language models used in intelligence, surveillance, and reconnaissance tasks. This leaderboard is not about general trivia but focuses on the reasoning, reporting, and restraint needed by analysts in sensitive scenarios. The goal? To identify which models can be trusted with critical intelligence work.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

In a recent evaluation conducted on July 17, 2026, 14 models were tested across 300 specialized tasks. The results are publicly available, but the actual test set remains secret, deliberately designed to prevent models from training on it. A private held-out set exists, and the difference between public and held-out scores reveals potential memorization, ensuring transparency about model generalization.

Leading the pack is Claude Fable 5, with a score of 67.77, earning it the top Band A ranking. Notably, a new entrant, Kimi K3 from Moonshot, has made an impressive debut at #3 with a score of 64.65, placing it above all GPT and Gemini models on the leaderboard. This newcomer is classified in Band B and shows promising potential in this specialized arena.

VigilSAR public LLM leaderboard
The leaderboard — compare bands, not rank numbers. Source: vigilsar.com/benchmark

The leaderboard also segments models into bands rather than precise ranks, with confidence intervals indicating the margin of error. The GPT-5.x family occupies Bands C-D, while Gemini models sit in Bands E-F. Interestingly, one locally-runnable open model has been scored as “sovereign-deployable”, meaning it could be deployed in real-world scenarios, highlighting the importance of practical deployment considerations in the evaluation.

The purpose behind this initiative is clear: the site emphasizes that “vendor claims are not evidence“. The operators built this evaluation to objectively measure which models are suitable for trust-based intelligence tasks, independent of vendor influence. They also publish confidence intervals, held-out gaps, a pinned reference row, and per-model economics, including cost-per-correct-answer, to promote transparency and honesty.

For the broader audience, this ranking offers a rare glimpse into the AI models that defense professionals might rely on in the future. The secretive nature of the test questions underscores the seriousness of the evaluation — models are tested on private, unseen tasks to ensure genuine capability. Interested readers can explore the current standings at the public leaderboard.

Powered by Thorsten Meyer AI


Amazon

AI model deployment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

defense AI analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

secure AI inference hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Jack Clark Says It Out Loud — Reading the Co-Founder’s 60%/2028 Estimate on Automated AI R&D

Anthropic co-founder Jack Clark publicly estimates a 60% probability that autonomous AI systems capable of self-advancement emerge by 2028.

Technology Is Never Neutral: Pope Leo XIV’s AI Encyclical, and the Empty Chairs in the Room

Pope Leo XIV’s first encyclical addresses AI’s moral challenges, highlighting the influence of those who develop and finance technology, with a focus on Anthropic’s role.