LLM surpasses pre-screening capabilities of traditional NLP and rule-based algorithms

Clinical-trial eligibility criteria are written in nuanced language, while database queries work best with fixed, predefined attributes.
How rule-based screening works
A trial prescreening process begins with the electronic health record. Clinicians document a patient’s diagnosis, treatments, laboratory results, imaging findings, and so on. Some of this information is already structured but much of the most relevant information, on average 80%, sits in free-text notes, reports, and correspondence.
Traditional NLP, entity-recognition, and rule-based algorithms attempt to identify facts in that free text and map them into a standard data model such as OMOP. A phrase such as “ER-positive breast cancer” may become a standardized condition or observation; a medication mention may become a drug-exposure record; and a laboratory result may become a measurement with a value and date.
In parallel, the trial’s inclusion and exclusion criteria are converted into a computable cohort definition. Creating this query often already requires a cooperation between a clinical trial coordinator that has the medical knowledge, and a more technical profile to understand query'ing and database models. In practice, building and refining such a query can be a lengthy and resource-intensive process. For more on the time this can take, see our previous blog post, where we evaluated the process in a real-life pilot at ZAS (ZAS pilot insights: the practical challenges of queries and the promise of GenAI). The result is a query often only covering a limited number of criteria of proxies to identify patients in the OMOP database. In the most strict sense it is simply filterering on a general age and diagnosis age or biomarker for example.
The quality of this list highly depends on whether the structured representation faithfully captures the full meaning of both the patient record and the protocol. In practice, we observe that these lists still contain many false positives.

The meaning and nuances get lost when simplifying medical data to pre-defined attributes
OMOP provides an excellent common structure for clinical data, but it does not automatically preserve the nuance of a note or a protocol. A rule may recognize “angina,” for example, without distinguishing active, historical, suspected, or explicitly ruled-out angina. The same problem arises when broad terms such as “ulcer” are mapped to a fixed concept without their clinical context.
Protocol criteria also depend on relationships that are difficult to reduce to isolated fields: a time window, an exclusion, a sequence of prior treatments, an exact biomarker variant, or compound AND/OR logic. SQL can encode these requirements, but only when every fact is available, normalized, dated, and linked correctly. Similar insights are published in the abstract of Adam Petranovich, clearly showing the greater amount of false positives that inevitably pop-up when using this methodology.
Why rule-based pre-screening falls short
Rule-based queries can be useful as a quick first filter, especially for clear, structured criteria such as age, sex, a documented diagnosis, or a laboratory threshold. But they are less reliable when eligibility depends on clinical nuance.
Structured attributes lose clinical context.
Query building is time-consuming.
A limited set of proxy criteria create false positives.
LLMs can understand what rules miss
Anyone who has used ChatGPT, Gemini, Claude, or a similar tool has experienced that technology has moved beyond simple keyword matching. Today’s LLMs can interpret context and distinguish active disease from a historical diagnosis, link a biomarker result to the relevant tumour type, reconstruct a sequence of prior treatments, and recognise when a condition has been explicitly ruled out. They can also assess more complex relationships involving negation, time windows, and exclusions.
Academic literature substantiates these claims. A multi-site real-world validation and an end-to-end study both reported high accuracy in criterion-level eligibility assessment. Monsana’s own validation studies show similarly strong results, and in some cases even stronger. We attribute this better performance in part to extensive criterion preprocessing and prompt optimisation tailored to each individual criterion, an additional layer of refinement that is not yet described in most published studies.
From Technical Capability to Clinical Impact

The next challenge is turning this capability into reliable day-to-day clinical impact. That requires security, integration into existing workflows, clear guardrails, transparent evidence for every assessment, scalability, and cost control. This is where Monsana works with hospitals: combining LLM-based clinical reasoning with a practical, human-in-the-loop service that reduces manual workload and helps teams avoid missing potential candidates in busy clinical practice. Adoption requires commitment and energy, but the technology has moved too far to ignore. The question is no longer whether LLMs can support trial pre-screening; it is whether organisations will adopt them early enough to benefit.
Ready to explore what this could mean for your trial-screening workflow? Talk to Monsana.



Comments