Comparative evaluation

Evaluating retrieval under changing evidence

Three oncology questions tested across five frontier models and Aprilio's research prototype. The comparison records omissions, outdated recommendations, and the source basis recovered by structure-first retrieval.

What are preferred first-line systemic treatment options for hepatocellular carcinoma?

What foundation models returned
Perplexity

Recommends camrelizumab + rivoceranib as a first-line option

FDA rejected twice (May 2024, March 2025)
DeepSeek

Includes sintilimab in preferred options

Unavailable in the United States
What the Aprilio prototype retrieved

For unresectable or metastatic HCC with preserved liver function, retrieves atezolizumab + bevacizumab, durvalumab + tremelimumab, and nivolumab + ipilimumab as preferred first-line options. Lenvatinib and sorafenib remain alternatives when immunotherapy or VEGF inhibition is unsuitable.

ASCO guideline + NCI PDQ + FDA approval, updated April 2025
Current approval surfaced: nivolumab + ipilimumab, April 11, 2025

The prototype separates preferred immunotherapy combinations from established alternatives and surfaces the approval date that changed the treatment landscape.

Why this happens

Training data goes stale. Oncology doesn’t wait.

Foundation models recall from training data. When that data is months old, the recommendations they surface can lag behind withdrawn drugs, new approvals, and updated guidelines.

the danger zone
Now
Gemini 3.1Jan 2025
Claude 4.6May 2025
GPT 5.4Sep 2025
DeepSeek V3.2Dec 2025
Zongertinib 1L approvedFeb 2026
Tazemetostat withdrawnMar 2026

Grounded Retrieval doesn’t guess. It reads the latest guidelines and tells you exactly where to verify.

Participate in external validation

We are inviting clinical researchers, domain experts, and governed knowledge-source partners to help test the method across new questions and source sets.

Built by a practicing medical oncologistValidated against 5 frontier modelsProvisional patent pending

Frequently asked questions

Reliability, repeated

One correct answer is not enough.

Aprilio repeats difficult questions to reveal variance that one-shot benchmarks hide.

Aprilio / internal test

20/20

stable

Detected the late-breaking update and returned exact provenance on every run.

Evaluated comparison lanes

0/20

update missed

Missed the newly approved regimen and returned inconsistent citation behavior.

Missing evidence triggers a halt, not a guess.

The system preserves a visible boundary when a required node or pathway is absent.

auditable halt state

Internal evaluation: one complex AML question repeated at temperature 0.3. External validation pending.