Aprilio / internal test
20/20
Detected the late-breaking update and returned exact provenance on every run.
Comparative evaluation
Three oncology questions tested across five frontier models and Aprilio's research prototype. The comparison records omissions, outdated recommendations, and the source basis recovered by structure-first retrieval.
“What are preferred first-line systemic treatment options for hepatocellular carcinoma?”
“Recommends camrelizumab + rivoceranib as a first-line option”
“Includes sintilimab in preferred options”
“For unresectable or metastatic HCC with preserved liver function, retrieves atezolizumab + bevacizumab, durvalumab + tremelimumab, and nivolumab + ipilimumab as preferred first-line options. Lenvatinib and sorafenib remain alternatives when immunotherapy or VEGF inhibition is unsuitable.”
The prototype separates preferred immunotherapy combinations from established alternatives and surfaces the approval date that changed the treatment landscape.
Why this happens
Foundation models recall from training data. When that data is months old, the recommendations they surface can lag behind withdrawn drugs, new approvals, and updated guidelines.
Grounded Retrieval doesn’t guess. It reads the latest guidelines and tells you exactly where to verify.
We are inviting clinical researchers, domain experts, and governed knowledge-source partners to help test the method across new questions and source sets.
Reliability, repeated
Aprilio repeats difficult questions to reveal variance that one-shot benchmarks hide.
Aprilio / internal test
20/20
Detected the late-breaking update and returned exact provenance on every run.
Evaluated comparison lanes
0/20
Missed the newly approved regimen and returned inconsistent citation behavior.
Missing evidence triggers a halt, not a guess.
The system preserves a visible boundary when a required node or pathway is absent.
Internal evaluation: one complex AML question repeated at temperature 0.3. External validation pending.