Explainer · Medical agent evaluation · 3 October 2026 · 7 minute read
OpenTumorBoard tests medical models against real cancer-care discussions
OpenTumorBoard pairs recorded specialist discussions with patient evidence and final decisions. It tests whether models can answer real clinical questions or simulate a whole meeting, revealing substantial gaps while providing examples for further training.
The short version
- The benchmark preserves actual specialist exchanges and board conclusions, rather than constructing reference answers from case descriptions.
- Its 611 patient cases support separate tests of individual specialist responses and complete simulated discussions.
- The best reported scores remain limited: 3.43 out of 5 for specialist answers and 2.78 for final-decision alignment.
- Training on recorded discussions improves a small model, but clinician validation of the scoring system remains unfinished.
Medical knowledge is not the same as a board decision
A multidisciplinary tumor board brings cancer specialists together to review a patient and decide what to do next. Their evidence includes images, molecular findings and the history of treatment responses. The task is not simply to name a plausible treatment. Specialists must reconcile different interpretations and identify constraints that change which options remain appropriate.
Limited specialist availability can delay these meetings and subsequent decisions. That motivates attempts to use language models as virtual specialists. But testing such systems requires more than checking whether they know medicine. An evaluation needs to capture the questions clinicians actually ask, their exchanges with colleagues and the decision that emerges.
The paper identifies several gaps in earlier benchmarks. Examination questions and questions written from case descriptions miss the distribution of questions raised in meetings. Interactive tests often involve a patient or medical record, rather than a specialist team. Simulated panels may also tackle cases without a recorded real discussion, leaving no corresponding team process to compare against.
Anqi Li and colleagues introduce OpenTumorBoard to supply that missing reference. Their contribution is a benchmark and training resource built from recorded meetings, not a demonstration that an automated board is ready to make clinical decisions.
Keep the evidence, the exchanges and the conclusion together
The central idea is to preserve a discussion trajectory: the ordered contributions through which specialists review evidence, propose actions and revise their views. Each case pairs that trajectory with a patient summary, presentation slides and the board's conclusion. Reference answers therefore come from observed exchanges, rather than being written retrospectively from a case description.
OpenTumorBoard contains 611 cases from 219 meetings on YouTube, covering 12,534 minutes of recordings. It includes 19,157 discussion turns involving ten specialist roles. For comparison, the authors identify MTBBench as containing 66 patient cases. The distinction is not only the case count: OpenTumorBoard retains the discussions associated with the decisions.
The authors define two tasks. Specialist Turn asks a model, assigned a particular specialist role, to answer a question from a recorded meeting. Board Simulation instead asks a model to generate the full exchange among specialists and its final conclusion. Both use the case summary and slides, but one tests a local contribution while the other tests the entire decision process.
That distinction matters when interpreting performance. Answering one question with the discussion context supplied is different from generating the sequence of perspectives needed to reach a board-level decision. Board Simulation uses a model to produce the multi-specialist discussion; it does not require separate models for each role.
Turn recordings into cases and testable decisions
The curation process starts with YouTube recordings whose titles mention tumor boards and whose duration exceeds twelve minutes. Transcript-based filtering removes generic lectures. The pipeline also excludes recordings that lack usable slides or do not follow the tumor board format. It then processes sound and images separately before joining them by time.
From meeting recordings to paired evaluation references
- Filter recordingsKeep patient-case meetings with usable slide footage.
- Process speechTranscribe utterances, identify voices, assign roles and separate cases.Process slidesFind and crop slides, remove near-duplicates, retain readable clinical content.
- Match timestampsConnect displayed evidence with the corresponding discussion.
- Extract paired recordsBuild case summaries, slide sets, specialist questions, responses and conclusions.
- Prepare evaluation inputsStandardize question wording and check cases against decision leakage.
For audio, WhisperX transcribes speech and pyannote assigns utterances to distinct voices. GPT-5.4 maps those voices to specialist roles and separates meetings into patient cases. The visual track identifies slides, crops their contents, removes near-duplicates and keeps readable clinical material. GPT-5.4 also captions the retained slides.
Timestamps connect each displayed slide to the corresponding speech. GPT-5.4 then extracts the pre-meeting summary and the conclusion reached during discussion. The pipeline retains patient-specific questions, filters out generic knowledge questions and standardizes their wording. The authors also describe a check against leaking the decision into the case input, with details deferred to an appendix.
A recorded prostate case illustrates why the sequence matters. An early proposal would repeat radiation to the prostate bed. The radiation oncologist identifies a previous dose of 71.8 Gy, and the proposal is withdrawn. After 32 turns, the board settles on surveillance guided by PSA doubling time. The example shows a specialist contribution changing an earlier plan, rather than merely adding another opinion.
Evaluation uses Qwen3.8-27B to score generated outputs against references. Clinical equivalence measures how faithfully an answer retains the recorded specialist's clinical meaning. Conclusion alignment measures agreement with the board's decisions about treatment, surgery, next actions and trial matching. Both use a scale from 1 to 5. Additional checks flag consequential clinical errors and claims lacking sufficient evidence.
The recordings are split into training, validation and test partitions using a 60/10/30 allocation. The test partition has 184 cases and 4,844 questions drawn from 66 recordings. Splitting recordings, rather than describing only a pool of questions, makes clear the level at which the source material is partitioned.
Real specialist questions expose substantial gaps
The authors report that DeepSeek-V4-Pro in reasoning mode leads Specialist Turn with clinical equivalence of 3.43 out of 5. Among the nine evaluated models, four score below 3. Across models, critical-error rates run from 5.6% to 43.5%, and unsupported-claim rates range from 10.0% to 74.5%. These checks expose problems beyond imperfect wording.
Medical specialization does not guarantee stronger responses in this evaluation: three of the four medical models occupy the bottom positions. Questions about supporting evidence are particularly difficult. For the strongest model, unsupported claims occur in 11.8% of responses overall, compared with 31.5% for trial suggestions and 23.5% for evidence discussion.
Board Simulation evaluates 14 models. The authors report that Gemini 3.7 Flash performs best, but its conclusion alignment reaches only 2.78 out of 5. This is a comparison with recorded board decisions, not with individual specialist answers. The scores describe different tasks and should not be treated as interchangeable measures.
The training experiment tests whether the recorded exchanges can help Qwen2.5-VL-3B. Supervised finetuning, training the model to reproduce example discussions and conclusions, uses 366 cases. Reinforcement learning then adjusts the trained model using rewards that include agreement between its conclusion and the board's decision.
Discussion-based training improves Qwen2.5-VL-3B alignment
Unit: Conclusion alignment, scored from 1 to 5
The authors report alignment of 1.58 for the base model, 1.67 after supervised finetuning and 1.86 after reinforcement learning. This controlled progression compares stages of the same model, unlike the ranking across different models. It supports the usefulness of discussion-based training, while the final score still leaves substantial disagreement with recorded decisions.
Reference quality and score validity remain separate issues
Three medical experts reviewed a random subset containing 15 cases. All reviewed cases received high ratings for coverage, factuality and fidelity to the board consensus. For extracted specialist answers, 92.6% received high correctness ratings and 94.4% received high ratings for support from the discussion. These findings support the sampled references, not an assumption of error-free extraction throughout the collection.
Crucially, reviewing extracted references does not validate the model-generated performance scores. The authors explicitly leave agreement between their automated judge and clinicians for future work. Clinical equivalence and conclusion alignment therefore remain rubric-based model judgments, rather than established clinician assessments of the generated answers.
The public-recording source also limits representativeness. These meetings may not reflect the full range of institutional practice. The authors state that they did not independently screen for patient identifiers beyond protections already applied to the public recordings. Public availability should not be mistaken for an additional privacy review by the benchmark team.
Other conclusions need similar restraint. Alignment measures agreement with a recorded board, not patient outcomes. The reported association between longer discussions and better alignment across models does not establish that forcing longer discussions improves decisions. And the training experiment on Qwen2.5-VL-3B alone does not establish that the gains transfer to other models.
A more grounded target for medical-agent research
Researchers evaluating medical agents gain a way to test case-specific reasoning against observed specialist behavior. Engineers developing simulated boards gain examples of how evidence review and competing perspectives lead to a decision. The separate tasks also distinguish a weak individual response from a failure to assemble an entire discussion.
The practical lesson is narrower than clinical readiness. General reasoning ability and medical specialization do not, by themselves, ensure faithful decisions on this benchmark. Recorded discussions offer a useful training target, and the reported gains justify further investigation. Stronger claims will require clinician-calibrated scoring and evidence beyond agreement with the recorded consensus.
Terms used here
- Multidisciplinary tumor board
- A meeting where cancer specialists review a patient together and decide on next steps.
- Discussion trajectory
- The sequence of specialist contributions, evidence reviews and revisions leading to a board conclusion.
- Specialist Turn
- The benchmark task of answering a recorded board question in an assigned specialist role.
- Board Simulation
- The benchmark task of generating an entire specialist discussion and its final conclusion.
- Clinical equivalence
- How faithfully a generated answer preserves the clinical meaning of the recorded specialist response.
- Conclusion alignment
- How closely a generated conclusion agrees with the decisions of the recorded board.
- Supervised finetuning
- Training a model to reproduce example outputs, here recorded discussions and conclusions.
- Reinforcement learning
- Training that adjusts a model using rewards, here including agreement with board decisions.
The work
OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
- Authors
- Anqi Li, Zhixuan Ge, Yixuan Duan, Jiarong Qian, Chi-Yu Chen, MingYu Lu, Huan-Yu Hsu, Yu Gu and 4 more
- Published
- 26 September 2026
- Code
- github.com
- Project page
- huggingface.co
Related reading
This explainer was written by AI from the source text and checked against it. Read the source for the full detail. How this site works