Validation & Benchmarks
How VetGeni validates its clinical answers
VetGeni is powered by Wiley-licensed veterinary references and a live corpus of peer-reviewed clinical guidelines. This page shows exactly how we test what the AI says — and what we do not claim.
29
live guideline sources
Peer-reviewed, versioned, inventoried
3
evaluation lanes
Knowledge, reasoning, citations
80% / 75% / 90%
benchmark-gate floors
Minimum MCQ / reasoning pass / citation coverage
What VetBench tests
VetBench is VetGeni’s in-house benchmark for guideline-backed clinical answering. Three lanes, each with a different failure mode in mind. Question sets are kept private to protect their integrity — only aggregate results are ever published.
Knowledge (MCQ)
Original multiple-choice questions written specifically for VetBench — never recalled or paraphrased from NAVLE or any other licensing exam — covering topics such as vaccination protocols, kidney-disease staging, cardiology staging, antimicrobial stewardship, and CPR guidelines. Scored deterministically against a fixed answer key.
Scored — automated
Clinical reasoning (rubric)
Open-ended clinical questions graded against veterinarian-written rubrics by a separate AI judge — a different model from the one that writes the answers, disclosed with each run below. Grades the judge is not confident in are held for manual review and excluded from the pass rate.
Scored — automated
Citation coverage
Checks that the sources cited alongside each answer actually match the guideline organizations and DOIs the answer should rest on — attribution is tested, not assumed.
Scored — automated
Benchmark gate: a scored run must reach at least 80% MCQ pass rate, 75% clinical-reasoning pass rate, and 90% citation coverage, or the run fails. These floors are pinned by automated tests and can only hold or rise — never be quietly lowered.
Current results
Run public-run-v3 on 2026-08-18 · code version 2449e7b8a0
| Lane | Cases | Result | Floor |
|---|---|---|---|
| Knowledge (MCQ) | 12 | 100% correct | 80% |
| Clinical reasoning (rubric) | 4 | 75% pass | 75% |
| Citation coverage | 3 | 100% covered | 90% |
The live guideline corpus
Every guideline document in VetGeni’s retrieval corpus, listed by bibliographic record. Guideline text itself is never republished here or in answers — citations are bibliographic only.
| Guideline | Org | Year | DOI | License |
|---|---|---|---|---|
| 2024 AAHA Fluid Therapy Guidelines for Dogs and Cats | AAHA | 2024 | — | Free to read |
| 2023 AAHA Selected Endocrinopathies of Dogs and Cats Guidelines | AAHA | 2023 | — | Free to read |
| 2022 AAFP/AAHA Antimicrobial Stewardship Guidelines | AAHA | 2022 | — | Free to read |
| 2022 AAHA Pain Management Guidelines for Dogs and Cats | AAHA | 2022 | — | Free to read |
| 2021 AAHA Nutrition and Weight Management Guidelines for Dogs and Cats | AAHA | 2021 | — | Free to read |
| 2020 AAHA Anesthesia and Monitoring Guidelines for Dogs and Cats | AAHA | 2020 | — | Free to read |
| 2019 AAHA Dental Care Guidelines for Dogs and Cats | AAHA | 2019 | — | Free to read |
| 2018 AAHA Infection Control, Prevention, and Biosecurity Guidelines | AAHA | 2018 | — | Free to read |
| ACVIM-endorsed statement: consensus statement and systematic review on guidelines for the diagnosis and treatment of chronic inflammatory enteropathy in dogs | ACVIM | 2026 | 10.1093/jvimsj/aalaf017 | Open access |
| ACVIM consensus statement on the diagnosis of immune thrombocytopenia in dogs and cats | ACVIM | 2024 | 10.1111/jvim.16996 | Open access |
| ACVIM Consensus Statement on the management of status epilepticus and cluster seizures in dogs and cats | ACVIM | 2024 | 10.1111/jvim.16928 | Open access |
| ACVIM consensus statement on the treatment of immune thrombocytopenia in dogs and cats | ACVIM | 2024 | 10.1111/jvim.17079 | Open access |
| Updated ACVIM consensus statement on leptospirosis in dogs | ACVIM | 2023 | 10.1111/jvim.16903 | Open access |
| ACVIM consensus statement on pancreatitis in cats | ACVIM | 2021 | 10.1111/jvim.16053 | Open access |
| ACVIM consensus guidelines for the diagnosis and treatment of myxomatous mitral valve disease in dogs | ACVIM | 2019 | 10.1111/jvim.15488 | Open access |
| ACVIM consensus statement on the diagnosis and treatment of chronic hepatitis in dogs | ACVIM | 2019 | 10.1111/jvim.15467 | Open access |
| ACVIM consensus statement on the diagnosis of immune-mediated hemolytic anemia in dogs and cats | ACVIM | 2019 | 10.1111/jvim.15441 | Open access |
| ACVIM consensus statement on the treatment of immune-mediated hemolytic anemia in dogs | ACVIM | 2019 | 10.1111/jvim.15463 | Open access |
| ACVIM consensus statement: Guidelines for the identification, evaluation, and management of systemic hypertension in dogs and cats | ACVIM | 2018 | 10.1111/jvim.15331 | Open access |
| ACVIM consensus statement: Support for rational administration of gastrointestinal protectants to dogs and cats | ACVIM | 2018 | 10.1111/jvim.15337 | Open access |
| ACVIM consensus update on Lyme borreliosis in dogs and cats | ACVIM | 2018 | 10.1111/jvim.15085 | Open access |
| 2015 ACVIM Small Animal Consensus Statement on Seizure Management in Dogs | ACVIM | 2016 | 10.1111/jvim.13841 | Open access |
| ACVIM Small Animal Consensus Recommendations on the Treatment and Prevention of Uroliths in Dogs and Cats | ACVIM | 2016 | 10.1111/jvim.14559 | Open access |
| Antimicrobial use guidelines for canine pyoderma by the International Society for Companion Animal Infectious Diseases (ISCAID) | ISCAID | 2025 | 10.1111/vde.13342 | Open access |
| International Society for Companion Animal Infectious Diseases (ISCAID) guidelines for the diagnosis and management of bacterial urinary tract infections in dogs and cats | ISCAID | 2019 | 10.1016/j.tvjl.2019.02.008 | Open access |
| Antimicrobial use Guidelines for Treatment of Respiratory Tract Disease in Dogs and Cats: Antimicrobial Guidelines Working Group of the International Society for Companion Animal Infectious Diseases | ISCAID | 2017 | 10.1111/jvim.14627 | Open access |
| 2024 RECOVER Guidelines: Advanced Life Support. Evidence and knowledge gap analysis with treatment recommendations for small animal CPR | RECOVER | 2024 | 10.1111/vec.13389 | Open access |
| 2024 RECOVER Guidelines: Basic Life Support. Evidence and knowledge gap analysis with treatment recommendations for small animal CPR | RECOVER | 2024 | 10.1111/vec.13387 | Open access |
| 2022 WSAVA guidelines for the recognition, assessment and treatment of pain | WSAVA | 2022 | 10.1111/jsap.13566 | Free to read |
Documents still in license review are excluded from the live corpus until cleared — being useful is not a license.
How answers are verified
Benchmarks test the system from the outside. These safeguards run inside it, on every clinical answer.
Drug doses come only from verified tools
The voice copilot never invents a dose, dosing math, or a contraindication. Doses it speaks come from one place: a verified answer computed by the clinical runtime against the patient chart — relayed exactly as written, units spelled out in full, never rounded or extrapolated. If a verified answer cannot be fetched, it says so instead of guessing.
Independent AI verification passes
Clinical answers on the verified runtime are checked by a separate AI verifier before delivery, and voice answers are all-or-nothing: if a verified answer cannot be produced, the copilot declines rather than guesses. Clinical voice turns — drug doses among them — additionally face a challenger pass that actively argues against the draft answer before it is relayed. Where the chat dock falls back to an unverified legacy path, that answer carries no verification markers in its trust rail.
Sources cited with every knowledge answer — bibliographic only
When VetGeni’s clinical AI (Olivia) answers from the guideline corpus, the answer carries its sources: title, organization, year, DOI, and license. Reference text itself is never republished — attribution is bibliographic by design, and that constraint is enforced by automated tests.
Every stage is flight-recorded
A flight recorder logs each stage of an answer — request, verification, delivery, and failures — so problems are visible and auditable rather than silent. Fail-quiet behavior is treated as a bug.
The profession’s own checklist
In 2026, an ACVIM statement on artificial intelligence in veterinary medicine introduced a 12-item AI Validation Factor checklist for appraising AI tools — and advised that if the first two questions fail, a tool should not be used no matter how strong the rest looks. Here is our item-by-item self-assessment, including the items we have not yet met. The wording paraphrases the checklist; the authors’ exact questions are in the cited statement.
1.Does the tool answer the precise question you are asking?
Yours to ask — we discloseEach VetGeni capability states its job narrowly: draft documentation for veterinarian review, guideline-backed answers with bibliographic citations, dose math relayed only from the verified runtime, and simulated cases for training. VetGeni does not diagnose, and questions outside the corpus get a plain "outside my sources" rather than an improvised answer.
2.Does your patient resemble the population the tool was built on?
Yours to ask — we discloseEvery knowledge answer carries its sources — organization, year, DOI — so you can see exactly which guideline an answer rests on and judge its fit to the patient in front of you. The full corpus is listed above; documents still in license review never reach it.
3.Can you supply the correct, complete input the tool needs to be reliable?
Yours to ask — we discloseVerified answers are computed against the patient chart. When the chart lacks what an answer needs — a current weight for a dose, for example — the runtime declines and names what is missing instead of guessing.
4.Has functional testing taken place?
MetHundreds of automated tests gate every release, and the benchmark floors on this page are pinned by tests that can hold or rise but never quietly fall.
5.Has cross-validation been performed?
Not applicableCross-validation appraises a model you trained. VetGeni deliberately trains no clinical model of its own — answers come from general-purpose frontier models grounded in the licensed corpus — so we evaluate the assembled system with VetBench instead.
6.Has A/B testing taken place?
MetNew grading and answering configurations run in staged shadow evaluations alongside the current production configuration, and are promoted only when they hold or improve results.
7.Has ground-truth validation taken place?
MetThe knowledge lane is scored deterministically against fixed answer keys, reasoning rubrics are veterinarian-written, and documentation pipelines are probed against veterinarian hand-corrected gold cases before prompt changes ship.
8.Has performance, robustness, and usability testing taken place?
In progress — gap statedPerformance and robustness are tested by VetBench and pinned regression suites, and every production answer stage is flight-recorded. Usability is shaped by daily in-clinic use; formal usability studies have not yet been run.
9.Have ethics and bias testing taken place?
In progress — gap statedThe safeguards are tested — verifier and challenger passes, the no-citation-no-clinical-claim rule, and refusal behavior all have automated coverage. A formal bias audit has not yet been conducted, and we state that here rather than imply one.
10.Have cybersecurity and privacy been addressed?
MetDocumented plainly on the Security & Data Ownership page, including what we are not: VetGeni holds no company-level SOC 2; our infrastructure providers maintain theirs. Higher-education security documentation is available to institutions on request.
11.Has the tool been tested in the real world, by independent parties, with ongoing quality assurance?
In progress — gap statedVetGeni runs in real clinical use daily, the results above come from a pinned, re-runnable benchmark, and production incidents feed a regression suite so the same failure cannot ship twice. Independent third-party evaluation has not yet happened — we publish our methodology and invite it.
12.Is the tool — and your use of it — compliant with applicable rules?
Yours to ask — we discloseNo dedicated regulator certifies veterinary clinical AI today. VetGeni aligns with the data-protection and education-records obligations that apply to it, and the practice-side half of this question — local veterinary rules on records, delegation, and telemedicine — stays with the licensed veterinarian, which is where the statement places it.
Framework: Bertin F–R, Lawrence J, Niessen SJM, Pinard CJ, Reagan KL, Rentko V. “The influence, promise, and potential perils of artificial intelligence in veterinary medicine: a call for improved awareness and literacy.” J Vet Intern Med 2026;40(1). doi.org/10.1093/jvimsj/aalaf045. This is VetGeni’s self-assessment — ACVIM and the statement authors have not reviewed or endorsed VetGeni.
A broader independent index of veterinary-AI position statements is the Veterinary AI Policy Tracker maintained by Dr. Candice Chu, DVM, PhD, DACVP (Texas A&M University).
What we do not claim
- VetGeni is not SOC 2 certified as a company. Our infrastructure providers (AWS, Supabase, Vercel) each maintain SOC 2 Type 2 audits of their platforms — the security page states the full posture plainly.
- AI answers are not a substitute for professional veterinary judgment. VetGeni supports clinical decisions; it does not make them.
- Benchmark items are original. Nothing in VetBench is recalled or paraphrased from NAVLE or any other licensing examination.
- AI-generated output requires professional review. Every answer and every document is a draft for a licensed veterinarian to review, edit, and approve.
The full security and compliance posture lives on the Security & Data Ownership page.
Judge It on Your Own Cases
Start your free 14-day trial. No credit card required. Every AI answer arrives with its sources.
About this page
Author
VetGeni Clinical Content Team
Veterinary Content Team
VetGeni
Reviewed by
VetGeni Clinical Review Team
Medical Review Board
VetGeni