Keiji AI LogoKeiji AI

For real-world evidence & epidemiology

AI agents for real-world evidence

Most analyst time goes to cleaning and joining data, and every new question restarts the engineering. TrialMind turns a feasibility question into an answer in minutes — across the sources you already license, running where the data lives.

Weeks → minutes

Feasibility turnaround at a precision oncology diagnostics company

~3×

analyst throughput, with roughly 40% less dependency on external analytics

Zero

raw records leaving your environment — the control plane operates on metadata

What we hear from RWE leaders

The bar moved. The capacity did not.

Answers arrive after the decision

Feasibility takes weeks of data engineering before any analysis starts. Claims, EHR, and registry data do not join cleanly, and each new question restarts the work.

Regulatory scrutiny, not internal

FDA wants the protocol and analysis plan before the analyses run, an audit trail from data extraction onward, and source data available for inspection.

The remit keeps widening

Feasibility counts, then comparative effectiveness where no trial exists, safety commitments with dates attached, burden of illness. Same molecule, same data, same four people.

Sources: FDA's Real-World Evidence Program · Considerations for the Use of RWD and RWE (final guidance, August 2023) · RWE: Considerations Regarding Non-Interventional Studies (draft guidance, March 2024)

The remit, covered

Seven workflows we will work on with you

Feasibility is where most teams start, not where the remit ends. The same cohort, the same definitions, and the same audit trail carry through all seven.

01

Feasibility with an attrition funnel

When it comes up

Clinical asks how many patients like this exist, with an unrealistic deadline

What you get

The count, the funnel showing how each criterion moved it, and the code

02

Comparative effectiveness by target trial emulation

When it comes up

A comparative question no trial will answer, or a single-arm study needing an external control

What you get

The emulated protocol stated, then executed: propensity or IPTW adjustment, balance diagnostics, and the sensitivity analysis a reviewer will ask for

03

Treatment patterns and drug utilization

When it comes up

Post-launch — who is getting this, in what sequence, and instead of what

What you get

Adoption, switching, sequencing, and adherence by line and biomarker. In oncology, line-of-therapy reconstruction with every rule made explicit

04

Burden of illness and unmet need

When it comes up

Medical affairs needs the disease characterized before anything can be argued

What you get

Incidence, prevalence, outcomes gaps, and utilization — one picture instead of five commissioned studies

05

Safety signals and post-marketing commitments

When it comes up

A signal appears, or a PMR/PMC has a filed protocol and a fixed date

What you get

Longitudinal and disproportionality analyses, and committed studies run to the agreed protocol with the audit trail it requires

06

Eligibility stress-test and generalizability

When it comes up

Protocol design before criteria are fixed, or a Diversity Action Plan

What you get

Per-criterion cohort impact, and your planned population against the real-world one by stratum

07

Standing cohort refreshed on each data cut

When it comes up

Quarterly refresh of a definition already agreed

What you get

The same definition re-run, with a diff against the prior cut

Where we start: feasibility. Speed is the headline, but the funnel is what makes it credible — a number without one is just a faster guess.

The second one matters most. An external control arm is target trial emulation at its highest stakes — the method that defends a single-arm submission also answers the comparative questions you will never run a trial for.

First, a data-fitness question: can the sources you license support the claim, which variables are missing or proxied, and what limitation survives. Better found in week one than in review.

See it working

One workflow, end to end

HER2-altered metastatic NSCLC, from cohort definition to survival comparison in one session. Real screenshots, not a mock.

Step 1

Define the cohort and the question

Specify the population — here, metastatic NSCLC with HER2 alterations — and the objective. The attrition table comes back with counts and percentages at every step, so the cohort can be inspected before anything is analyzed.

TrialMind attrition table for metastatic NSCLC patients with HER2 alterations

Step 2

Run the survival analysis

Time-to-event analyses — here Kaplan–Meier overall survival — run across the cohorts with consistent definitions of index date, censoring, and follow-up. Plan, code, and output are all retained.

Kaplan-Meier overall survival curves by HER2 treatment and testing cohort

Deployment & governance

Built to pass privacy and data governance review

Two things decide whether an RWE project happens: where the data has to go, and whether the method survives scrutiny. Both are answered in the architecture, not in a policy commitment.

Secure by design

The AI control plane is separate from the data plane. Raw records never leave your environment — an architectural property, not a promise. SOC 2 report, certificates, and completed vendor questionnaires available on request.

Regulated-grade access and audit

Access is enforced at the reasoning layer, not the interface: the agent cannot reason over data the user may not see, whatever a prompt asks. Immutable audit trail, electronic signature to 21 CFR Part 11, ALCOA+ throughout.

Data-neutral by design

Works across the claims, EHR, registry, and genomic sources you already license — no new purchase, no vendor decision. Fits the estate in place: OMOP and your own schemas, alongside SAS, R, Python, Spotfire, and Tableau.

Models built on clinical research

Domain-specific models outperform general-purpose ones on clinical tasks, and the comparison is published rather than asserted. Test it against a cohort you have already built — that comparison is the pitch.

SOC 2 Type IIISO 27001HIPAA compliant · BAAs in placeRead our security and compliance details →

Published evidence

The methods, mapped to the papers

Read them before trusting anything downstream. The two to send first are DSWizard, the clearest statement of the division of labour we propose, and the MLHC reconstruction paper, the one a methodologist will want to argue with.

What it supportsPaper
Agent-run analysis you can checkDSWizard — reliable multi-step data-science agents for biomedical research (Nature Biomedical Engineering, 2026) · Can Large Language Models Replace Data Scientists in Clinical Research? (Nature Biomedical Engineering, 2024)
Longitudinal EHR modelingHALO — synthesis of high-dimensional longitudinal records (Nature Communications, 2023) · PromptEHR (EMNLP, 2022) · SynTEG (JAMIA, 2021)
Prediction on real-world populationsMediTab — EHR pre-training for zero-shot prediction on trial populations (IJCAI, 2024) · Evidence-driven spatiotemporal hospitalization prediction (Nature Communications, 2023)
Comparators without patient-level accessIndividual patient data reconstruction from published trials — accepted at Machine Learning for Healthcare (MLHC) 2026
Population representativenessGenerative balancing for equity in medical machine learning (npj Digital Medicine, 2025)
Patient–cohort matchingTrialGPT (Nature Communications, 2024) — deployed at the NIH with a reported 40% average time saving

See all publications →

FAQ

The questions procurement asks first

Do we need to buy data from you?

No. TrialMind is data-neutral and works across the claims, EHR, registry, and genomic sources you already license — no additional purchase, and no need to settle which vendor wins internally.

Does raw patient data leave our environment?

No. TrialMind deploys into your environment and analyses run where the data already lives. Raw records never move, which is what makes privacy and governance review passable.

Can the output support a regulatory submission?

That is what the audit trail is for. Cohort definitions, index dates, censoring rules, and analysis steps are captured and reproducible, and the methods are published so the approach can be defended by citation. Formal computer-system validation is a separate track, and we would rather tell you where that stands than let the controls imply it.

Who actually runs it — an epidemiologist or a data engineer?

The epidemiologist who has the question. Natural-language cohort extraction is the point: it removes the data engineering cycle that currently sits between the question and the answer. The SQL, R, Python, and cohort definitions your team has already validated import as they are.

Why not just use a general-purpose model?

That option is in the room whether or not anyone names it. The difference shows up when the analysis has to be reproduced, defended to a regulator, or run again next quarter on refreshed data — and when access has to be enforced at the reasoning layer rather than at the interface.

How does it handle joining claims and EHR data?

TrialMind reasons across sources rather than requiring them pre-harmonised into one warehouse. Where your data already sits in a warehouse, we connect to it rather than move it.

The first step

Bring one question you have already answered

You know the right answer. We run it and you check our work — the funnel and the code, not just the count. Minutes versus weeks is the demo.

Start here

One feasibility question, answered the slow way

We run it, you check our work: the count, the attrition funnel, and the generated code.

Then

The comparative question you could not answer

Where no trial exists and the analysis stalled on confounding. We show the emulated protocol and diagnostics before anyone commits — so you judge a method, not a demo.