Portrait of Carl Luis PöhlAvailable from January 2027 · Munich, open to relocation

Hi, I'm Carl. I build AI that doctors can check.

I'm in my final year of medical school in Dresden and finishing a PhD on AI for cancer tumor boards. I also studied computer science, and I co-founded Halverd, an AI consultancy. Most of my work comes back to one question: how do you know a clinical AI system is actually right?

The box above only knows what's on this page, and every answer links to where it came from. If something isn't here, it tells you instead of guessing. That's the same rule I hold my research to.

Where I've studied and worked

University Hospital DresdenDKFZ HeidelbergKather Lab · EKFZUniversity Hospital HeidelbergTU DresdenRWTH AachenEPFLHKUST

96.6%

guideline adherence at the prostate-cancer tumor board

4

university hospitals in the multi-site validation

44 × 500

configurations × held-out real cases benchmarked

0.936

external AUROC on TCGA-PAAD frozen slides

01

Ventures

What I'm building right now.

06/2026 – present

Halverd

Co-founder & AI Consultant · AI consultancy for professional services

  • Shipped multiple AI systems into production, spanning candidate matching, team planning and assistants inside the messaging channels people already use, each returning output with the reasoning attached.
  • Built the agent infrastructure for a client deployment and the evaluation framework behind it, with test sets and scoring so output quality stays measurable after release.
  • Co-delivered a £15,000+ recruitment-sector build, integrated with the workflow software their consultants already use.
AgentsEvalsProduction

12/2025 – present

ConcordAInce

Co-founder · Tumor-board decision-support AI

  • Translating the PhD tumor-board decision-support work toward clinical deployment, with privacy-preserving infrastructure that keeps patient data within the hospital perimeter.
  • Dresden Exists LifeTech incubator. Y Combinator Spring 2026 interview.
OncologyOn-premStartup

03/2026 – present

Healicus

Co-founder · Natural-medicine AI assistant

  • Consumer-facing RAG reference for natural medicine, with each entry grounded in a Cochrane review, EMA HMPC monograph, EFSA authorised health claim or major-journal RCT.
  • Personalised pre-generation checks for drug interactions, contraindications and allergies, with severity-tiered alerts and auditable event records.
RAGSafetyConsumer

02

Research

My PhD, and the projects that grew around it.

01/2025 – present

Multi-agent decision support for the prostate-cancer tumor board

PhD Candidate (Dr. rer. medic.), AI in Medicine · Dept. of Urology, University Hospital Dresden

Supervised by Sherif Mehralivand and Marcus Sondermann, in cooperation with Jakob Kather and Lars Hilgers (Kather Lab, EKFZ).

  • Developed hybrid multi-agent tumor-board decision support combining structured patient data with EAU and German S3 oncology guidelines.
  • Retrospective validation at the prostate-cancer tumor board: 96.6% guideline adherence, with additional guideline-compliant therapy options identified in 57.8% of cases. External validation across four university hospitals, including a prospective arm, ongoing.
  • Introduced guideline conformity as a separate evaluation endpoint after finding that the fine-tuned system agreed with tumor-board decisions in 99.2% of cases yet ranked below every structured design on conformance.
  • Compared 44 configurations × 500 held-out real cases across six architectures. Validated the LLM judge against eight board-certified raters, each scoring the same 50 cases, and designed their scoring rubric and adjudication protocol. Multi-agent, tree and hybrid designs were most guideline-conformant on intermediate- and high-risk cases, while zero-shot, RAG and fine-tuned approaches performed better on easy cases.
5 more
  • Study leadership. secured participation from Heidelberg, Mannheim, Zurich and Basel for the multi-site study, persuading two initially sceptical full professors without formal authority. Wrote every ethics submission, reworking it through two rounds of committee feedback to approval. Secured the project's HPC compute allocation on a written proposal.
  • Architecture. combined verified LLM extraction, deterministic decision trees, multi-agent escalation, case retrieval and PubMed evidence. The published benchmark evaluates an earlier configuration of this system.
  • Data. benchmarked on 1,000 tumor-board cases drawn at random from an 11,488-case institutional archive, split 500 test and 500 development, with fine-tuning on 2,521 separate archive cases and zero patient-level overlap. Held-out cases were excluded from retrieval, and retrieval was never combined with fine-tuned models.
  • Deployment. integrated the open-source system into the hospital infrastructure, running on-premise in the clinic on commodity hardware so no patient data leaves the hospital. GDPR-compliant by design.
  • Safety. built automated evaluation and CI/CD checks for guideline adherence, evidence traceability, hallucinations and contraindications, with cite-or-abstain behaviour and clinician oversight. Red-teamed clinical failure modes and built regression tests using constitutional-AI-style synthetic cases.
Multi-agent LLMsLLM-as-judgeClinical validation

12/2025 – present

Pancreatic histogenomics

DKFZ Heidelberg + University Hospital Heidelberg, Surgery

In collaboration with Andrea Bauer.

  • Benchmarked seven foundation-model encoders on 527 fresh-frozen pancreatic whole-slide images, achieving 96.8% balanced accuracy on the 475-slide PDAC/pancreatitis/normal subset in patient-level cross-validation.
  • Evaluated calibrated abstention on the five-class differential: 95.1% accuracy at 77% coverage (confidence ≥0.90). External PDAC-vs-rest detection achieved 0.936 AUROC on 257 TCGA-PAAD frozen slides.
  • Evaluated pathology-transcriptomics fusion on 303 slides from 301 patients. Rare cystic and neuroendocrine tumours received greater genomic weighting across four random seeds.
1 more
  • Developed an inflammatory-infiltrate model on 311 pancreatic slides: 83.2% balanced accuracy and 0.897 macro-AUROC in patient-grouped cross-validation. Its predictions tracked nine immune-gene readouts more closely on average than pathologist grades (mean Spearman ρ 0.535 vs. 0.462 across 203 matched slides).
Pathology FMsCalibrationMultimodal

04/2025 – present

Calibrated Parkinson’s progression-subtype prediction

MD Thesis (Dr. med.) · Dept. of Neurology, University Hospital Dresden

Supervised by Tom Hähnel.

  • Developed calibrated Parkinson’s progression-subtype prediction from 17 routine scores in 409 PPMI patients, using per-patient OLS slopes and intercepts.
  • Compared Random Forest, XGBoost and logistic regression with missing-score handling and class-conditional conformal abstention.
  • Assessed fixed-model predictions against motor, cognitive and biomarker outcomes in the later PPMI 2.0 cohort. Independent subtype-level validation remains outstanding.
Conformal predictionNeurology

04/2024 – 11/2024

DNA language models across 19 cancer types

Student AI Researcher · Kather Lab (EKFZ / TU Dresden)

  • Benchmarked Hyena-DNA, DNABERT-2 and NucleotideTransformer-v2 across 19 TCGA cancer types (7,674 patients). Hyena-DNA embeddings with a residual feed-forward network achieved 0.94 mean-class AUROC at 54.96% overall accuracy.
  • Improved NucleotideTransformer-v2 accuracy by almost 10% with mutation-centred windows, approaching DNABERT-2. Results suggested greater sequence-shift sensitivity with k-mer than byte-pair tokenisation.
Genomic FMs

03

Open source

Code and products you can try yourself.

cite-or-abstain

Python · MIT

Clinical LLM evaluation harness for citation verification, unsupported confidence and abstention, with human-validated LLM judging, DeepEval integration and reproducible audit reports.

EvalsDeepEval

promotionshub

TypeScript · Next.js · live

Platform where medical students find doctoral positions and read verified reviews of supervision, live with 135 positions across 28 research groups. Ratings are structurally separated from paid listings so promotion cannot influence them, named scores stay hidden until four independent verified reviews corroborate them, and verification documents are discarded after checking.

ProductTrust & safety

parkinson-subtype-predictor

Python · MIT · live demo

Live Streamlit app with calibrated predictions, abstention, SHAP and counterfactual explanations. Single-patient and batch workflows, plus reproducible PPMI analysis respecting the data-use agreement. Demo: parkinson-subtype-predictor.onrender.com.

StreamlitExplainability

hf-titration-assistant

FastAPI + React/TypeScript · MIT

Heart-failure prototype combining explainable deterioration risk with guideline-based GDMT titration. XGBoost/LightGBM on 98 trajectory features. Zigong evaluation with patient-level splits and bootstrap confidence intervals.

CardiologyXGBoost

04

Manuscripts & talks

Papers on the way, and where I've presented the work.

Manuscripts in preparation

First-author throughout. Four manuscripts complete, in co-author review.

  • PROMPT. a deterministic guideline tree achieves similar automated guideline-conformance scores to multi-agent LLMs for prostate cancer tumour-board decisions. Pöhl, Sondermann, Kather, Mehralivand.
  • Pan-pancreatic diagnosis on fresh-frozen sections: a foundation-model characterization of the open multi-class differential, calibrated abstention, and class-conditional multimodality. Pöhl et al. DKFZ Heidelberg + University Hospital Heidelberg.
  • A foundation model reads a molecularly anchored inflammatory-infiltrate axis from fresh-frozen pancreatic histology. Pöhl et al. University Hospital Heidelberg + DKFZ Heidelberg + University of Verona.
  • A Calibrated, Abstaining Prediction Model for Parkinson’s Disease Progression Subtypes from a Variable Set of Routine Clinical Scores. Pöhl, Falkenburger, Hähnel. TU Dresden. Basis of the Dr. med. thesis.

Selected talks

  • Nov 2026Poster (accepted), ESMO AI & Digital Oncology Congress, Berlin: multi-agent decision support for the prostate-cancer tumor board.
  • Sep 2026Invited lab meeting, Dewey Lab, University of Cambridge: “Agreeing with the tumor board is not the same as being right.”
  • Jul 2026Invited talk, DKFZ Heidelberg: prostate-cancer decision support.
  • Jun 2026Invited talk, Urology, University Hospital Basel: guideline-grounded AI.
  • May 2026Invited lab meeting, Krauthammer Lab, University of Zurich: clinical LLMs.
  • Mar 2026Poster, ELSA TrustworthyAI4Health: prostate-cancer tumor boards.
  • Oct 2024Invited talk, Connectome Fall Symposium: DNA-language models.

05

Industry

Two internships, at Bayer in Tokyo and EY in Munich.

08/2022 – 09/2022

EY, Munich

Intern · Transaction, Strategy & Execution

  • Co-developed the kick-off strategy for a client's market exit and prepared execution workshops with global stakeholders.

02/2020 – 03/2020

Bayer, Tokyo

Data Science Intern · Crop Science Division

  • Built the division's first data-driven targeting model for sales-activity planning, with no in-house data science function to build on, replacing gut-feel prioritisation. A retrospective backtest indicated savings of up to ¥120M. Presented findings to the head of the Asia-Pacific division.

06

Education

Medicine and computer science, mostly at the same time.

2025 – expected Q1 2027

Dr. rer. medic. + Dr. med.

TU Dresden

  • Dr. rer. medic. (PhD equivalent), AI in Medicine: multi-agent decision support for the prostate-cancer tumor board.
  • Dr. med. (medical doctorate, research thesis): calibrated progression-subtype prediction in Parkinson’s disease.

09/2020 – present

Medical studies (State Examination)

TU Dresden · Semmelweis University, Budapest

  • Second national medical examination (M2) passed 2025, good (2.0), ~top 20% nationally. Final examination December 2026.
  • At TU Dresden since 10/2022, after pre-clinical studies at Semmelweis University, Budapest (09/2020–07/2022). Physiology and Biochemistry examinations at the top grade.
  • International clinical rotations across six settings: cardiology (Buenos Aires) · surgery + dermatology (Windhoek, Namibia) · visceral surgery + neurosurgery (University Hospital Heidelberg) · gastroenterology + haematology/oncology (University Hospital Zurich) · radiology (TUM University Hospital, Munich) · general practice (southern Germany).

2017 – 2024

B.Sc. Computer Science

RWTH Aachen · EPFL · HKUST

  • RWTH Aachen (10/2018–11/2024): final grade good (2.0), completed part-time alongside medicine. Mathematics very good (1.3), minor in Business Administration.
  • EPFL Lausanne (09/2019–03/2020): exchange semester in Computer Science, SEMP scholarship.
  • HKUST Hong Kong (09/2017–10/2018): Dual Degree Program in Technology & Management (Computer Science & Business). GPA 3.74/4.0, Dean's List both semesters.
2 more
  • Abitur, German School Tokyo Yokohama (05/2017): final grade 1.6, Mathematics and Physics as written subjects. Youngest graduate, at age 16.
  • Grading scales

    German 1.0 best, 4.0 passHKUST 4.0 best.

Awards, scholarships & programmes

  • McKinsey Capstone Programme, member (2020 – present).
  • Dresden Exists LifeTech Incubator Program, incubator place for ConcordAInce (2025).
  • Porsche IT Scholarship, competitive technology scholarship (2020–2021).
3 more
  • BASF Science Prize (Chemistry) · German Physical Society Prize (Physics), Abitur 2017.
  • SEMP Scholarship, Swiss-European Mobility Programme, EPFL exchange (2019–2020).
  • HKUST Dean's List, top 10% of cohort (2017–2018).

Teaching & academic engagement

  • Tutor, Biochemistry, Faculty of Medicine, TU Dresden (04/2023–06/2025): review and exam-preparation sessions for cohorts of 100+ medical students.
  • Executive Committee Member, Connectome Neuroscience Society, TU Dresden (10/2023–10/2025): organised interdisciplinary neuroscience and neurosurgery seminars and workshops.

07

Toolkit

What I work with, and what I do when I'm not working.

Technical skills

  • Core ML

    PyTorchNumPypandasscikit-learnXGBooststatsmodels
  • LLMs & training

    LoRA / QLoRAmulti-agent orchestrationprompt engineeringstructured outputsAnthropic / OpenAI SDKsvLLM
  • Evaluation & safety

    InspectDeepEvalOpenAI Evals adaptersLLM-as-judgered-teamingregression testingevidence attributioncalibrationprompt-injection mitigation
  • Biological foundation models

    GigaPathUNI / UNI2Virchow2CONCHCTransPathTITANABMIL / CLAM-MB / TransMIL / DTFD-MILgated cross-modal fusiongenomic sequence models
  • Retrieval & RAG

    FAISSChromapgvectorLangChainLlamaIndexhybrid retrieval
  • Backend & deployment

    FastAPINext.jsPostgreSQLSupabaseAWSHetznerDockerCI/CDGitHub ActionsStreamlitpytestGiton-premise deployment

Languages & interests

  • German, Spanish and English (native) · French (fluent) · Italian (conversational) · Japanese and Mandarin (basic).
  • Raised across Ecuador, Mexico, Spain, Germany, Japan and Norway.
  • Concert piano performance · chess · alpine skiing · sailing.

Contact

Working on clinical AI? I'd like to hear about it.

Email is the best way to reach me.