ISCB GMDS 2026: AI, Statistical Rigor, and Lessons for RPACT Code

Four days in Freiburg on AI guardrails, reproducible statistical code, adaptive trials, and the responsibility that remains with us.
Conferences
AI
RPACT Code
rpact
Author

Friedrich Pahlke

Published

October 1, 2026

From 28 September to 1 October, I attended ISCB GMDS 2026 in Freiburg im Breisgau. With more than 1,300 participants on site and up to 12 parallel sessions, choosing a personal programme was essential. Mine focused on AI, particularly its relevance to our RPACT Code project, alongside adaptive clinical trial design.

These are the ideas I brought back for colleagues working with statistical software: where AI helps, where it fails, and what those failures teach us about building useful tools.

Conference participants in Freiburg's Konzerthaus, facing the stage and presentation screen.

The Konzerthaus during Tuesday’s keynote programme.

Monday, 28 September: Making AI useful and its results dependable

Day in brief: Health data infrastructure, practical AI applications, and adaptive trials shared a common requirement: the surrounding workflow matters as much as the method.

Opening and digital health

After the Opening Session, Cathie Sudlow delivered “Future of Population Health Research: a UK Health Data Perspective.” She argued for treating health data as national infrastructure, with secure access, linkage across sources, and sustained public engagement. Reusable data-curation pipelines deserve investment and recognition: they make subsequent research possible.

The Digital Health Tools and Engagement session brought this down to implementation. Contributions from H Räther and LS Brandstetter linked adoption to stakeholder involvement, usability, feedback, and support. In “MeLTSy: A MeSH-based system for linking medical terminology with laypeople synonyms,” C Sautier showed both the usefulness and the limits of terminology mapping: a plausible synonym can still be wrong in context.

R Roller, in “AI-Supported Chat System to Assist Telemedicine Staff in Monitoring Kidney Transplant Patients,” described message suggestions that staff reviewed before sending. Routine reminders worked better than messages about abnormal values; delayed data access and integration across hospitals remained practical obstacles.

The closest connection to RPACT Code

A highlight was “AI-based statistical tools with guardrails: building and validating a chatbot for sample size calculation” (E Carr, S Obi, OR Olaniran, D Shamsutdinova, F Zimmer, S Markham, D Stahl, G Forbes).

The team built a conversational interface to an R package for prediction-model sample size calculations, using a cache of precomputed scenarios. Initially, the LLM controlled input collection, tool calls, and the answer. It invented inputs, confused parameters, and even silently changed numerical results returned by the tool.

Three guardrails changed the architecture:

  • Deterministic input handling: extract and check parameters, reject unsupported values, and disclose approximations to cached scenarios.
  • Deterministic result display: show the package result and inputs directly, outside the LLM’s generated answer.
  • Curated statistical guidance: support explanations with a knowledge base, while recognising that instructions alone cannot guarantee compliance.

The input and output controls made the largest difference. In the evaluated final configuration, displayed sample sizes came unchanged from the package. Evaluation also checked whether the chatbot appropriately stopped when information was missing. User testing was still outstanding, so this was evidence about the tested conversations, not a guarantee for every real-world interaction.

For me, the valuable distinction was between helping someone understand a calculation and controlling the calculation itself.

D Dunkler’s “From Plan to Code: A Randomized Coding Experiment on How Specificity of a Statistical Analysis Plan Shapes Reproducibility” complemented this. Analysts and AI agents translated plans into code with substantial variation. More detailed plans generally reduced that variation; ambiguity and coding errors still mattered. Executable code does not automatically mean a faithful implementation.

LLMs for clinical workflows

The LLMs for Clinical Workflows session offered several useful checks on enthusiasm:

  • T Žvirblis, “Rule-Based Prompting for Extraction of TNM Stage and Tumor Grade from Population-Level Lithuanian Pathology Reports Using a Medical Large Language Model”: explicit extraction rules improved results, while missing evidence required an empty answer rather than an invented value.
  • C Demus, “LLM-Based Semi-Automated Information Extraction from Histopathological Reports in T-Cell Lymphoma Patients”: selecting relevant passages with rules before querying the model reduced hallucinations and computational effort. Manual verification remained part of the workflow.
  • F Alickovic, “From Free-Text to ICD-10 Codes: Comparison of Embedding Models and Retrieval-Augmented Generation for Automated Coding of German Tumor Diagnoses”: retrieving a correct candidate did not ensure the LLM selected it. Strong embedding baselines remained competitive, especially for exact codes.
  • J Sam, “Benchmarking AI Vibe Coding for Clinical Statistical Analysis: A Structured Evaluation in Pulmonary Hypertension Research”: ambiguous variable names produced convincing but incorrectly grouped summaries; another output omitted requested interaction terms. The small, single-run benchmark illustrated failure modes rather than establishing a model ranking.
  • V Vishnevskaya, “LLM-Assisted Generation of Off-Label Medication Requests in Oncology: A QUEST-Based Evaluation Framework”: fluent patient-history paragraphs could distort chronology or clinical relationships. Evaluation must look beyond readability.
  • “IRIS: A Multi-Agent Large Language Model Framework for Conversational Clinical Trial Navigation and Protocol Design” (A Abootalebi, P Messina, M Torchia, L Emili) connected registry retrieval to computational models. P Messina presented the work: language models helped retrieve and explain, while the simulator performed the calculations. Parts of the framework remained under development.

Master protocols and rare diseases

The invited session Innovations in Master Protocols Evaluating Multiple Treatments and (Rare) Diseases closed my scientific programme. JJ Lee, in “The successes and challenges of Bayesian adaptive platform trials in practice – lessons learned,” distinguished a functioning platform from a successful drug and emphasised simulation and operational readiness.

J Wason’s “Basket Trials for Efficient Precision Medicine Trials” explored information sharing across conditions, including rare diseases, with efficiency gains weighed against bias. SS Villar’s “Implementing Response-Adaptive Randomization in Rare-Disease Trials: Design Challenges and Exact Solutions” showed how discrete allocation updates could fit operational systems and stabilise adaptation; inference must account for the allocation design.

Personal photograph in front of Freiburg Minster.

A moment outside the conference rooms, by Freiburg Minster.

My RPACT takeaway. For RPACT Code, I would translate these lessons into explicit assumptions, checked tool inputs, unchanged numerical outputs, and tests of both correctness and completeness. For rpact, simulation and transparent design specifications remain central. A conversational interface should make those foundations easier to use and inspect.

Tuesday, 29 September: Human judgment, external evidence, and AI ethics

Day in brief: AI competence needs domain competence. Likewise, additional data only strengthen evidence when their assumptions and limitations are understood.

Two complementary keynotes

In “The End of Data-Driven Medicine: Entering the Era of AI-Native Science,” Marylyn Ritchie argued for developing AI skills alongside scientific expertise. Her phrase “human in the lead” captured the difference between retaining decision-making authority and merely reviewing automated decisions. Training must preserve the ability to reason, interpret, and work when AI is unavailable or wrong.

Annette Kopp-Schneider’s “Leveraging External Data in Clinical Trials: Principles, Perils & Potential” examined what information borrowing really buys. In settings with a uniformly most powerful test, apparent power gains can disappear when borrowing and non-borrowing methods are compared at the same type I error level. Benefits require defensible assumptions about how external and current data relate. Robust borrowing is not a universal safeguard; assumptions, tuning parameters, and operating characteristics need transparent reporting.

Ethical and technical evaluation

The workshop On the Ethical and Technical Aspects for Deployment and Evaluation of AI-Driven Clinical Decision Support Systems, chaired by Stefan Rühlicke and Zully Ritter, connected technical metrics with their consequences for people.

In “The Fairy Tale of Fair Clinical AI,” R Roller stressed subgroup and intersectional evaluation: a strong overall score can hide poor performance for particular patients. RE Martín-Peña’s “The Uncoded Patient: Structural Blind Spots in Clinical AI Evaluation” moved the question upstream, to information lost when patient experience becomes a variable.

S Moazemi discussed feature-level agreement with clinical expertise; SF Rashid connected performance, calibration, and domain shift with ethical consequences. S Trivedi, in “Navigating the Ethical Landscape of Artificial Intelligence in Clinical Applications: The Role of Algorithmic Impact Assessment Tool,” considered assessment throughout development and deployment, rather than a one-time checklist.

The afternoon Signature Session: AI and Me offered a different perspective: Jonas Gerigk improvised on double bass with SOMAX2, followed by discussion. Human–AI interaction became something to hear and experience.

A double bassist on the Konzerthaus stage beneath the ISCB GMDS 2026 conference screen.

Jonas Gerigk and SOMAX2 in the signature session.

Friedrich Pahlke (left) and Lukas Widmer from Novartis (right) at a conference social gathering.

With Lukas Widmer from Novartis (right), with whom I regularly exchanged ideas throughout the conference.

My RPACT takeaway. RPACT Code should help users understand assumptions and assess alternatives. Its evaluation should cover ambiguous requests, unsupported tasks, and different levels of statistical experience. External-data discussions reinforced the same principle for trial planning: report the conditions behind a result, not just the attractive number.

Wednesday, 30 September: Adaptive decisions and statistical literacy

Day in brief: Dose-finding research showed how assumptions shape decisions. The afternoon connected statistical literacy with critical use of AI, before Andre Dekker examined what it takes to make AI useful in clinical practice.

Dose Finding and Adaptive Trial Design

In the morning oral session, LA Widmer’s “Design Informed Prior Calibration for BLRM and Emax Models in Oncology Phase I Dose Escalation Trials” showed that seemingly diffuse priors can imply implausible dose-response curves and stall escalation. Priors should be checked against plausible curves and early patient outcomes.

I Vannier, “Comparisons of intra-patient dose escalation approaches in phase I clinical trials,” found that performance depended strongly on assumptions about toxicity across cycles. No approach was best across all settings.

Y MEI’s “Personalized Bayesian Phase I/II Trial Design Incorporating In-vitro Data: Application to Organoid and Cystic Fibrosis” investigated linking organoid information to clinical responses while carrying uncertainty through the model. A Vuorinen’s “PK/PD-Integrated Bayesian Platform Design for Phase II Regimen Optimisation” used exposure and activity information to inform safety and efficacy decisions across dosing schedules.

The contribution “Using GPC with multiple prioritized outcomes to analyze dose-ranging trials” (S Pandevant, M Buyse, S Chevret, L Biard) highlighted a subtle limitation: when efficacy dominates pairwise comparisons, safety may contribute little to the combined result.

C Voller’s “Optimal response-adaptive randomisation using Bayesian decision theory in group sequential designs” supplied a useful benchmark: in the studied scenarios, early stopping drove the largest reduction in expected loss, and a well-chosen fixed allocation remained competitive.

From left to right: Anika Grosshennig, Michael Brendel, Silke Szymczak, and Friedrich Pahlke, former colleagues from IMBS Lübeck.

Reunited with former colleagues from IMBS at the University of Lübeck, where we worked on our doctorates together. From left to right: Anika Grosshennig, Michael Brendel, Silke Szymczak, and me.

Teaching statistical literacy

In the afternoon I attended Teaching Statistical Literacy in Times of Political Changes, chaired by Ursula Berger and Carolin Herrmann. A question raised in the introduction captured the challenge: why learn biostatistics when AI can produce an analysis? The five contributions gave substantial reasons to keep learning.

  • WJ Radermacher, “Data ethics: the missing puzzle piece in applied statistics education,” distinguished individual, institutional, and societal responsibilities. Ethical practice needs support at all three levels. He also described an early idea for using an AI tutor to help students recognise ethical conflicts in practical statistical work.
  • K Ickstadt, “Rethinking statistics education at universities,” described interdisciplinary teaching in which statistics and data science students work with students from other subjects. Data quality, ethics, interpretation, and communication belong alongside software skills. Her examples of tandem teaching made this concrete: students learn to explain and question algorithms with people who bring different expertise.
  • S Hoffmann, presenting “Teaching responsible research practices that improve the communication and understanding of evidence and uncertainty” (S Hoffmann, AL Boulesteix), questioned how much uncertainty our intervals actually capture. Assumptions about measurement error, missing data, model specification, and generalisability can leave important uncertainty unrepresented. She advocated interpretable summaries, such as predicted risks, and teaching examples that expose multiple defensible analyses rather than always ending with one significant result. Preregistration helps address selective reporting, but does not remove analytical variability.
  • G Rauch, “Democracy under attack -Empirical research and statistical inference in times of fake news,” argued that students should learn to turn everyday claims into research questions and assess the evidence behind them. AI can generate code quickly; judging sources, assumptions, and the relevance of the question still requires independent thought. An authoritative answer, whether from a person or an AI tool, deserves scrutiny.
  • EC Wit, “Nature of Statistical Evidence: Lessons from Covid-19,” encouraged asking what a number is being compared with and which assumptions produced it. His graph example showed how changing the time axis could change the apparent story without changing the data. Other examples concerned the interpretation of life expectancy and the consequences of ignoring population heterogeneity in epidemic models.

The panel discussion returned to questions that precede any statistical test: what is the research question, the endpoint, and the target population? An audience contribution proposed using AI as a questioning partner in research training, helping students examine their plans. This was a planned teaching approach, not an evaluated success. The discussion also resisted equating trustworthiness with certainty: evidence can be trustworthy while its limitations remain explicit.

AI for Better Health Care

Andre Dekker’s keynote, “AI for Better Health Care,” connected AI development with the practical experience of an oncology clinic. His Personal Health Train approach brings analyses to data held at participating institutions. This requires interoperable data and technical and organisational governance, as well as learning algorithms.

He distinguished AI used to save work from AI used to improve treatment decisions. Both require evaluation in the setting where they will be used. A COVID imaging model that performed well in its original hospital failed when transferred to Maastricht, illustrating the consequences of differences in scanners and patient populations. Federated learning across more diverse centres offered one route towards broader validation and better generalisation.

His implementation examples were equally relevant. Technically good outputs may not be accepted in routine care, and automating easy tasks can leave staff with a more demanding workload. He challenged the assumption that checking every AI output necessarily delivers the intended efficiency gain, while showing that humans and AI make different mistakes. For me, the key question was how to design and evaluate their collaboration. Retrospective associations, he also stressed, still need further testing before they justify causal treatment claims.

My RPACT takeaway. For rpact users, the morning reinforced the value of comparing designs under multiple scenarios and inspecting what drives a decision. For RPACT Code, I see value in asking clarifying questions before generating code, making assumptions visible, and helping users explore sensitivity analyses. Dekker’s examples add a practical evaluation question: does assistance reduce the total effort of producing a checked analysis, including review and correction? These are development considerations, not claims that the specialised methods presented are already available in our products.

Thursday, 1 October: Trustworthy AI requires accountable researchers

Day in brief: My final morning connected hallucination detection with the broader responsibility for statistical evidence.

The Early Career Researcher Day Session 2 included J Miebs’ “Hallucination Detection in Large Language Models.” He distinguished factual errors from failures to follow the requested task or supplied context. That distinction matters for statistical assistance: an answer can sound plausible while addressing a different question.

He compared three approaches to detecting questionable output:

  • Internal uncertainty scores use token probabilities or entropy, sometimes giving more weight to words that carry the meaning of an answer. They require access to model information that may be unavailable through a given service.
  • Repeated-answer comparisons examine consistency or group answers by meaning. These can be used with less access to the model, but require additional generations.
  • Source-based checks extract individual claims and compare them with external evidence.

The discussion made an important limitation explicit: an uncertainty score is not a calibrated probability that an answer is wrong, and uncertainty does not equal hallucination. A high score can prompt another check; a consistent answer still needs evidence. Miebs also highlighted the risk of an early error propagating through a chain of agents. For RPACT Code, I see these methods as possible additional signals for review, alongside deterministic checks of inputs, generated code, and numerical results.

In “Trustworthy Biostatistics in the Age of AI: Research Integrity as Professional Practice,” U Mansmann placed responsibility along the entire chain from question and assumptions to analysis and interpretation. AI can help structure reasoning, write code, and suggest alternatives, but it cannot take responsibility for their validity.

A useful distinction concerned documentation: even when an AI answer cannot be reproduced exactly, we should be able to reconstruct how it influenced the scientific work and what was checked. Mansmann also urged researchers to retain the practice of thinking, calculating, and writing themselves, preserving the skills needed to judge AI output.

Slide contrasting AI assistance with the statistician's responsibility for correctness, confidentiality, reproducibility, citation accuracy, and interpretation.

U Mansmann’s slide on the statistician’s continuing responsibilities.

My RPACT takeaway. I came away with a concrete direction for RPACT Code: make the path from request to code to checked result visible, and help the statistician remain in control. Freiburg strengthened my interest in AI precisely because the conference paired working examples with careful accounts of their limitations. That combination is what I want to carry into our development work.

Photographs from my conference collection.