Skip to main content

Validation & Benchmarks

How VetGeni validates its clinical answers

VetGeni is powered by Wiley-licensed veterinary references and a live corpus of peer-reviewed clinical guidelines. This page shows exactly how we test what the AI says — and what we do not claim.

29

live guideline sources

Peer-reviewed, versioned, inventoried

3

evaluation lanes

Knowledge, reasoning, citations

80% / 75% / 90%

benchmark-gate floors

Minimum MCQ / reasoning pass / citation coverage

What VetBench tests

VetBench is VetGeni’s in-house benchmark for guideline-backed clinical answering. Three lanes, each with a different failure mode in mind. Question sets are kept private to protect their integrity — only aggregate results are ever published.

Knowledge (MCQ)

Original multiple-choice questions written specifically for VetBench — never recalled or paraphrased from NAVLE or any other licensing exam — covering topics such as vaccination protocols, kidney-disease staging, cardiology staging, antimicrobial stewardship, and CPR guidelines. Scored deterministically against a fixed answer key.

Scored — automated

Clinical reasoning (rubric)

Open-ended clinical questions graded against veterinarian-written rubrics by a separate AI judge — a different model from the one that writes the answers, disclosed with each run below. Grades the judge is not confident in are held for manual review and excluded from the pass rate.

Scored — automated

Citation coverage

Checks that the sources cited alongside each answer actually match the guideline organizations and DOIs the answer should rest on — attribution is tested, not assumed.

Scored — automated

Benchmark gate: a scored run must reach at least 80% MCQ pass rate, 75% clinical-reasoning pass rate, and 90% citation coverage, or the run fails. These floors are pinned by automated tests and can only hold or rise — never be quietly lowered.

Current results

Run public-run-v3 on 2026-08-18 · code version 2449e7b8a0

Release gate: passed
LaneCasesResultFloor
Knowledge (MCQ)12100% correct80%
Clinical reasoning (rubric)475% pass75%
Citation coverage3100% covered90%
Answering model gpt-5.6-sol · verifier gpt-5.5 · judge gpt-5.5 — the judge is a separate model from the one that writes the answers, and the full results artifact is committed in our repository.

The live guideline corpus

Every guideline document in VetGeni’s retrieval corpus, listed by bibliographic record. Guideline text itself is never republished here or in answers — citations are bibliographic only.

GuidelineOrgYearDOILicense
2024 AAHA Fluid Therapy Guidelines for Dogs and CatsAAHA2024Free to read
2023 AAHA Selected Endocrinopathies of Dogs and Cats GuidelinesAAHA2023Free to read
2022 AAFP/AAHA Antimicrobial Stewardship GuidelinesAAHA2022Free to read
2022 AAHA Pain Management Guidelines for Dogs and CatsAAHA2022Free to read
2021 AAHA Nutrition and Weight Management Guidelines for Dogs and CatsAAHA2021Free to read
2020 AAHA Anesthesia and Monitoring Guidelines for Dogs and CatsAAHA2020Free to read
2019 AAHA Dental Care Guidelines for Dogs and CatsAAHA2019Free to read
2018 AAHA Infection Control, Prevention, and Biosecurity GuidelinesAAHA2018Free to read
ACVIM-endorsed statement: consensus statement and systematic review on guidelines for the diagnosis and treatment of chronic inflammatory enteropathy in dogsACVIM202610.1093/jvimsj/aalaf017Open access
ACVIM consensus statement on the diagnosis of immune thrombocytopenia in dogs and catsACVIM202410.1111/jvim.16996Open access
ACVIM Consensus Statement on the management of status epilepticus and cluster seizures in dogs and catsACVIM202410.1111/jvim.16928Open access
ACVIM consensus statement on the treatment of immune thrombocytopenia in dogs and catsACVIM202410.1111/jvim.17079Open access
Updated ACVIM consensus statement on leptospirosis in dogsACVIM202310.1111/jvim.16903Open access
ACVIM consensus statement on pancreatitis in catsACVIM202110.1111/jvim.16053Open access
ACVIM consensus guidelines for the diagnosis and treatment of myxomatous mitral valve disease in dogsACVIM201910.1111/jvim.15488Open access
ACVIM consensus statement on the diagnosis and treatment of chronic hepatitis in dogsACVIM201910.1111/jvim.15467Open access
ACVIM consensus statement on the diagnosis of immune-mediated hemolytic anemia in dogs and catsACVIM201910.1111/jvim.15441Open access
ACVIM consensus statement on the treatment of immune-mediated hemolytic anemia in dogsACVIM201910.1111/jvim.15463Open access
ACVIM consensus statement: Guidelines for the identification, evaluation, and management of systemic hypertension in dogs and catsACVIM201810.1111/jvim.15331Open access
ACVIM consensus statement: Support for rational administration of gastrointestinal protectants to dogs and catsACVIM201810.1111/jvim.15337Open access
ACVIM consensus update on Lyme borreliosis in dogs and catsACVIM201810.1111/jvim.15085Open access
2015 ACVIM Small Animal Consensus Statement on Seizure Management in DogsACVIM201610.1111/jvim.13841Open access
ACVIM Small Animal Consensus Recommendations on the Treatment and Prevention of Uroliths in Dogs and CatsACVIM201610.1111/jvim.14559Open access
Antimicrobial use guidelines for canine pyoderma by the International Society for Companion Animal Infectious Diseases (ISCAID)ISCAID202510.1111/vde.13342Open access
International Society for Companion Animal Infectious Diseases (ISCAID) guidelines for the diagnosis and management of bacterial urinary tract infections in dogs and catsISCAID201910.1016/j.tvjl.2019.02.008Open access
Antimicrobial use Guidelines for Treatment of Respiratory Tract Disease in Dogs and Cats: Antimicrobial Guidelines Working Group of the International Society for Companion Animal Infectious DiseasesISCAID201710.1111/jvim.14627Open access
2024 RECOVER Guidelines: Advanced Life Support. Evidence and knowledge gap analysis with treatment recommendations for small animal CPRRECOVER202410.1111/vec.13389Open access
2024 RECOVER Guidelines: Basic Life Support. Evidence and knowledge gap analysis with treatment recommendations for small animal CPRRECOVER202410.1111/vec.13387Open access
2022 WSAVA guidelines for the recognition, assessment and treatment of painWSAVA202210.1111/jsap.13566Free to read

Documents still in license review are excluded from the live corpus until cleared — being useful is not a license.

How answers are verified

Benchmarks test the system from the outside. These safeguards run inside it, on every clinical answer.

  1. Drug doses come only from verified tools

    The voice copilot never invents a dose, dosing math, or a contraindication. Doses it speaks come from one place: a verified answer computed by the clinical runtime against the patient chart — relayed exactly as written, units spelled out in full, never rounded or extrapolated. If a verified answer cannot be fetched, it says so instead of guessing.

  2. Independent AI verification passes

    Clinical answers on the verified runtime are checked by a separate AI verifier before delivery, and voice answers are all-or-nothing: if a verified answer cannot be produced, the copilot declines rather than guesses. Clinical voice turns — drug doses among them — additionally face a challenger pass that actively argues against the draft answer before it is relayed. Where the chat dock falls back to an unverified legacy path, that answer carries no verification markers in its trust rail.

  3. Sources cited with every knowledge answer — bibliographic only

    When VetGeni’s clinical AI (Olivia) answers from the guideline corpus, the answer carries its sources: title, organization, year, DOI, and license. Reference text itself is never republished — attribution is bibliographic by design, and that constraint is enforced by automated tests.

  4. Every stage is flight-recorded

    A flight recorder logs each stage of an answer — request, verification, delivery, and failures — so problems are visible and auditable rather than silent. Fail-quiet behavior is treated as a bug.

The profession’s own checklist

In 2026, an ACVIM statement on artificial intelligence in veterinary medicine introduced a 12-item AI Validation Factor checklist for appraising AI tools — and advised that if the first two questions fail, a tool should not be used no matter how strong the rest looks. Here is our item-by-item self-assessment, including the items we have not yet met. The wording paraphrases the checklist; the authors’ exact questions are in the cited statement.

  1. 1.Does the tool answer the precise question you are asking?

    Yours to ask — we disclose

    Each VetGeni capability states its job narrowly: draft documentation for veterinarian review, guideline-backed answers with bibliographic citations, dose math relayed only from the verified runtime, and simulated cases for training. VetGeni does not diagnose, and questions outside the corpus get a plain "outside my sources" rather than an improvised answer.

  2. 2.Does your patient resemble the population the tool was built on?

    Yours to ask — we disclose

    Every knowledge answer carries its sources — organization, year, DOI — so you can see exactly which guideline an answer rests on and judge its fit to the patient in front of you. The full corpus is listed above; documents still in license review never reach it.

  3. 3.Can you supply the correct, complete input the tool needs to be reliable?

    Yours to ask — we disclose

    Verified answers are computed against the patient chart. When the chart lacks what an answer needs — a current weight for a dose, for example — the runtime declines and names what is missing instead of guessing.

  4. 4.Has functional testing taken place?

    Met

    Hundreds of automated tests gate every release, and the benchmark floors on this page are pinned by tests that can hold or rise but never quietly fall.

  5. 5.Has cross-validation been performed?

    Not applicable

    Cross-validation appraises a model you trained. VetGeni deliberately trains no clinical model of its own — answers come from general-purpose frontier models grounded in the licensed corpus — so we evaluate the assembled system with VetBench instead.

  6. 6.Has A/B testing taken place?

    Met

    New grading and answering configurations run in staged shadow evaluations alongside the current production configuration, and are promoted only when they hold or improve results.

  7. 7.Has ground-truth validation taken place?

    Met

    The knowledge lane is scored deterministically against fixed answer keys, reasoning rubrics are veterinarian-written, and documentation pipelines are probed against veterinarian hand-corrected gold cases before prompt changes ship.

  8. 8.Has performance, robustness, and usability testing taken place?

    In progress — gap stated

    Performance and robustness are tested by VetBench and pinned regression suites, and every production answer stage is flight-recorded. Usability is shaped by daily in-clinic use; formal usability studies have not yet been run.

  9. 9.Have ethics and bias testing taken place?

    In progress — gap stated

    The safeguards are tested — verifier and challenger passes, the no-citation-no-clinical-claim rule, and refusal behavior all have automated coverage. A formal bias audit has not yet been conducted, and we state that here rather than imply one.

  10. 10.Have cybersecurity and privacy been addressed?

    Met

    Documented plainly on the Security & Data Ownership page, including what we are not: VetGeni holds no company-level SOC 2; our infrastructure providers maintain theirs. Higher-education security documentation is available to institutions on request.

  11. 11.Has the tool been tested in the real world, by independent parties, with ongoing quality assurance?

    In progress — gap stated

    VetGeni runs in real clinical use daily, the results above come from a pinned, re-runnable benchmark, and production incidents feed a regression suite so the same failure cannot ship twice. Independent third-party evaluation has not yet happened — we publish our methodology and invite it.

  12. 12.Is the tool — and your use of it — compliant with applicable rules?

    Yours to ask — we disclose

    No dedicated regulator certifies veterinary clinical AI today. VetGeni aligns with the data-protection and education-records obligations that apply to it, and the practice-side half of this question — local veterinary rules on records, delegation, and telemedicine — stays with the licensed veterinarian, which is where the statement places it.

Framework: Bertin F–R, Lawrence J, Niessen SJM, Pinard CJ, Reagan KL, Rentko V. “The influence, promise, and potential perils of artificial intelligence in veterinary medicine: a call for improved awareness and literacy.” J Vet Intern Med 2026;40(1). doi.org/10.1093/jvimsj/aalaf045. This is VetGeni’s self-assessment — ACVIM and the statement authors have not reviewed or endorsed VetGeni.

A broader independent index of veterinary-AI position statements is the Veterinary AI Policy Tracker maintained by Dr. Candice Chu, DVM, PhD, DACVP (Texas A&M University).

What we do not claim

  • VetGeni is not SOC 2 certified as a company. Our infrastructure providers (AWS, Supabase, Vercel) each maintain SOC 2 Type 2 audits of their platforms — the security page states the full posture plainly.
  • AI answers are not a substitute for professional veterinary judgment. VetGeni supports clinical decisions; it does not make them.
  • Benchmark items are original. Nothing in VetBench is recalled or paraphrased from NAVLE or any other licensing examination.
  • AI-generated output requires professional review. Every answer and every document is a draft for a licensed veterinarian to review, edit, and approve.

The full security and compliance posture lives on the Security & Data Ownership page.

Judge It on Your Own Cases

Start your free 14-day trial. No credit card required. Every AI answer arrives with its sources.