
Three Gates Between an ALS Biomarker and a Trial Endpoint
Most ALS biomarker conversations start with the wrong question. Which marker should we measure? The more useful question is what decision the measurement has to support, because a marker that is entirely adequate for one decision will fail at the next one without changing in any observable way. Selecting a blood biomarker for an ALS program is not a single choice made at the start. It is three sequential decisions, each with a different evidence bar, and programs stall at the transitions far more often than at the selection.
What does it take to move a blood biomarker from discovery to a trial endpoint in ALS?
Three distinct evidence packages. First, analytical validation adequate to the intended use. Second, evidence that a change in the marker reflects the drug’s mechanism rather than progression or peripheral physiology. Third, a formal qualification argument tied to a stated context of use. Each gate requires evidence the previous one did not.
Key takeaways
- Analytical validation is not one standard. Requirements scale with the decision the data will support, which is the core of the fit-for-purpose approach [4].
- Gate two is where tissue provenance becomes load-bearing. Attributing a change to a drug mechanism requires knowing the signal came from the tissue the drug acts on.
- Plasma NfL is measured with excellent precision and still varies with renal function [5], blood volume, and body mass index [6]. Precision and interpretability are separate properties.
- ALS has the field’s clearest precedent for gate three: tofersen received accelerated approval on the basis of reduced plasma NfL [1].
- Qualification is granted for a stated context of use, not for a marker in general [3]. A context-of-use statement written late is a context-of-use statement written badly.
Gate one: from measurable to fit for purpose
The first gate is analytical, and it is the one most programs think they have already cleared.
The confusion comes from treating validation as a single standard a method either meets or does not. It is not. The fit-for-purpose framework, which has governed biomarker method development for two decades, holds that validation requirements should scale with the intended use of the data and the regulatory weight that use carries [4]. An exploratory method supporting an internal go/no-go decision needs less than a method whose output will appear in a regulatory submission. The framework is iterative by design: you validate to the current decision, then revalidate upward when the decision changes.
What this means in practice is that a method adequate at gate one may be inadequate at gate three without a single reagent changing. Programs discover this late, usually when a regulator asks a question the validation package was never built to answer.
Four properties do most of the work at this gate. Precision, expressed as replicate variation across plates and across sites. Stability, meaning how the analyte behaves through the freeze-thaw and storage history the trial will actually impose on samples. Dilutional linearity, which tells you whether a result at one concentration is comparable to a result at another. And matrix effects, meaning whether the plasma itself interferes with the measurement in ways that differ between patients. A program that can state numbers for all four, in the matrix and the storage condition the trial will use, has cleared gate one. One that can state numbers for precision alone has cleared part of it and usually does not know which part.
Pre-analytical handling belongs in this gate rather than being treated as a downstream nuisance. When a standardization working group empirically tested variations in blood collection and handling across a panel of blood-based neurodegeneration markers, collection tube type alone produced different values for every marker assessed, and delayed centrifugation affected a subset [7]. For vesicle-based measurement the pre-analytical surface is larger still, and field guidance now covers pre-processing variables and separation requirements explicitly [8]. A multi-site ALS trial that has not fixed its handling protocol before first patient in has introduced a variance source it cannot model out afterwards. Reporting frameworks now exist for precisely this, developed in recognition that hundreds of pre-analytical protocols and more than forty distinct variables are in active use across blood vesicle research [12]. Recording which ones your trial used is not paperwork; it is what makes a result interpretable by someone who did not run it.
Figure 1. What each gate requires that the previous one did not
| Gate | Decision the marker supports | Evidence added at this gate |
|---|---|---|
| One: fit for purpose | Internal go/no-go; exploratory readout | Precision, stability, matrix effects, fixed pre-analytical protocol |
| Two: pharmacodynamic | Did the drug do what we designed it to do? | Attribution — that a change reflects mechanism, not progression or peripheral physiology |
| Three: endpoint | Regulatory decision on effectiveness | A qualification argument tied to a stated context of use, plus a link to clinical benefit |
Structured from the fit-for-purpose validation framework [4] and the FDA biomarker qualification evidentiary framework [3].
Gate two: from fit for purpose to pharmacodynamic
This is the gate that fails quietly, and the reason is attribution.
A pharmacodynamic biomarker has to answer a narrow question: did the drug engage its target and produce the intended molecular consequence? Answering it requires that a change in the measurement be attributable to the drug’s mechanism rather than to disease progression, to a change in the patient’s physiology, or to variation in a compartment the drug never touched. That is a different property from precision, and no amount of analytical performance supplies it. Plasma NfL makes the distinction concrete. It is a well-characterized marker of axonal injury and it is measured with excellent analytical precision. It also varies with renal function in older adults, an association that persists after adjusting for age, sex, and body mass index [5], and separately with blood volume and body mass index [6]. None of that is measurement error. It is real biology entering a total-plasma measurement from outside the compartment of interest. In a mechanism-of-action readout, a change of that origin is indistinguishable from the change you are trying to detect.
Timing is the second half of the attribution problem and gets less attention than it deserves. A pharmacodynamic readout only means something if the sampling schedule is aligned to the expected time course of the mechanism. Sample too early and the molecular consequence has not propagated. Sample too late and it has been absorbed into progression, at which point the readout is measuring disease course rather than drug action. For a mechanism with an unknown time course, the honest position at phase 1 is to sample densely and accept that you are characterizing the curve rather than testing a hypothesis about it. Programs that skip this step arrive at phase 2 with a schedule chosen for site convenience and a signal they cannot interpret.
The reproducibility literature in fluid biomarkers reaches the same conclusion from a different direction, identifying cohort composition, sample handling, and analytical variation as interacting sources of non-replication that have to be designed against rather than assessed afterwards [9]. The effect-size problem compounds it: biomarker associations reported in highly cited papers are consistently larger than the same associations in subsequent meta-analyses [10], so a pharmacodynamic effect size taken from the literature is probably optimistic.
What clears this gate is knowing where the analyte came from. Measuring a neuronal analyte inside a neuron-derived vesicle population rather than in whole plasma changes the denominator of the measurement, which is what makes attribution possible. Published work has demonstrated the approach for TDP-43 specifically, detecting elevated levels in plasma neuronal-derived vesicles in an amyloid-confirmed cohort where cerebrospinal fluid concentrations are too low for practical measurement [11]. The underlying measurement problem for TDP-43 is worth understanding separately; the point here is narrower, that provenance is the specific property gate two demands.
Gate three: from pharmacodynamic to endpoint
ALS has the clearest precedent in neurology for this transition, and it is worth reading carefully rather than triumphantly.
In the phase 3 VALOR trial, tofersen was compared with placebo over 24 weeks of dosing in adults with SOD1 ALS. The primary endpoint, change in the ALS Functional Rating Scale-Revised at week 28 in participants predicted to progress faster, was not met. The trial did show reductions in cerebrospinal fluid SOD1 as an indirect measure of target engagement, and reductions in plasma NfL [1]. Accelerated approval followed in April 2023 on the basis of the plasma NfL reduction, with a confirmatory trial in presymptomatic carriers as the condition.
Two things in that sequence matter for anyone designing a program now. The clinical endpoint failed and the biomarker carried the approval, which is the outcome biomarker strategy exists to produce. And the biomarker in question was not tissue-specific. Gate three was cleared on a marker that would struggle at gate two for a mechanism-specific readout, because what a regulator needed here was a link to clinical benefit in a defined population, not molecular attribution.
It is also worth being clear about what made that argument available. SOD1 ALS is genetically defined, which removes diagnostic ambiguity from the population. Plasma NfL had years of accumulated natural history data behind it, establishing how it behaves in untreated disease. And the condition’s severity and rarity limited the feasibility of running another placebo-controlled trial, which shifted the benefit-risk calculation. Remove any one of those and the same biomarker data supports a weaker case. A program without a genetically defined population or a natural history dataset is not one step behind the precedent; it is missing a precondition of it.
That distinction is what the qualification framework formalizes. Qualification is a conclusion that within a stated context of use, a marker can be relied on to have a specific interpretation in drug development [3]. The context of use is the operative unit: population, disease stage, the decision the marker informs, and the measurement method. A marker is never qualified in the abstract. The standardized vocabulary for these distinctions, including the separation between pharmacodynamic, prognostic, and surrogate roles, is set out in the joint FDA-NIH glossary [2], and using it precisely in early interactions with a regulator is cheaper than discovering later that your program and your reviewer meant different things by the same word.
Where ALS programs actually stall
ALS makes all three gates harder than Alzheimer’s does, for structural reasons rather than scientific ones.
Trials are short and populations are small, so there is less room to accumulate evidence across phases. Progression is heterogeneous, which widens the variance a pharmacodynamic signal has to clear. And the functional rating scale that anchors clinical assessment is coarse relative to the changes a mechanism-specific marker can detect over a short window.
There is also a compounding effect specific to small indications. In a large disease area, a marker accumulates evidence across many programs and sponsors, and any single program inherits that base. In ALS the base is thinner, which means individual programs carry more of the evidentiary burden themselves and get less benefit from work done elsewhere. That argues for generating more evidence earlier than the phase would seem to require, and for treating natural history data as an asset worth building rather than a cost to minimize.
The stalls follow a recognizable pattern. A program validates a method to exploratory standard, generates an encouraging phase 1 signal, and then finds at phase 2 that the method cannot support the decision it now needs to support. Or it builds a pharmacodynamic case on a marker that moves for reasons unrelated to the mechanism, and cannot separate drug effect from progression. Or it reaches end of phase 2 with good biomarker data and no context-of-use statement, having never specified in advance what claim the data was meant to support. The broader case for biology-based rather than symptom-based assessment in ALS is well established, as is the current therapeutic landscape these programs sit inside. The gap is procedural: knowing that blood biomarkers help does not tell you what evidence to generate, in what order.
Figure 2. Common stall points and what precedes them
| Stall | Where it surfaces | What was skipped earlier |
|---|---|---|
| Method inadequate for the decision | End of phase 1, or first regulatory interaction | Validating to the anticipated use rather than the current one |
| Drug effect inseparable from progression | Phase 2 interim | Establishing tissue provenance before building the mechanism argument |
| Pre-analytical variance across sites | Multi-site phase 2 | Fixing and documenting a handling protocol before first patient in |
| No qualification argument | End-of-phase-2 meeting | Drafting a context-of-use statement at program start |
What to build in advance
Four things are much cheaper to do early than to retrofit.
- Write the context-of-use statement first. Population, stage, the decision, the method. It will change, and having a version to revise is worth more than having none [3].
- Validate to the gate you are heading for, not the one you are standing in, wherever the cost difference is tolerable [4].
- Fix the pre-analytical protocol before the first site opens, and document it to the level current field guidance specifies [7][8].
- Decide early whether your program needs a tissue-specific readout. If the biomarker has to demonstrate mechanism, it does, and that decision drives the measurement approach rather than following from it. Platforms built for neuron-derived vesicle enrichment, including ExoSORT™, exist to serve this specific requirement.
The tofersen precedent is often read as evidence that blood biomarkers can now carry regulatory weight in ALS. That reading is correct but incomplete. What it established is that one marker, in one genetically defined population, with a decade of accumulated natural history behind it, could support one accelerated approval. Programs targeting mechanisms without that history will need to build the equivalent case themselves, gate by gate, and the ordering is not optional. The work now going into mechanism-specific blood measurement in ALS is aimed at the second gate in particular, which is where most programs currently lose the argument. NeuroDex builds biomarker programs around that sequence.
References
- Miller TM, Cudkowicz ME, Genge A, et al. Trial of antisense oligonucleotide tofersen for SOD1 ALS. New England Journal of Medicine. 2022;387(12):1099–1110. doi:10.1056/NEJMoa2204705
- FDA-NIH Biomarker Working Group. BEST (Biomarkers, EndpointS, and other Tools) Resource. Silver Spring (MD): Food and Drug Administration; Bethesda (MD): National Institutes of Health; 2016–. https://www.ncbi.nlm.nih.gov/books/NBK326791/
- U.S. Food and Drug Administration, Center for Drug Evaluation and Research and Center for Biologics Evaluation and Research. Biomarker Qualification: Evidentiary Framework. Draft Guidance for Industry and FDA Staff. December 2018. https://www.fda.gov/media/119271/download
- Lee JW, Devanarayan V, Barrett YC, et al. Fit-for-purpose method development and validation for successful biomarker measurement. Pharmaceutical Research. 2006;23(2):312–328. doi:10.1007/s11095-005-9045-3
- Akamine S, Marutani N, Kanayama D, et al. Renal function is associated with blood neurofilament light chain level in older adults. Scientific Reports. 2020;10(1):20350. doi:10.1038/s41598-020-76990-7
- Manouchehrinia A, Piehl F, Hillert J, et al. Confounding effect of blood volume and body mass index on blood neurofilament light chain levels. Annals of Clinical and Translational Neurology. 2020;7(1):139–143. doi:10.1002/acn3.50972
- Verberk IMW, Misdorp EO, Koelewijn J, et al. Characterization of pre-analytical sample handling effects on a panel of Alzheimer’s disease-related blood-based biomarkers: Results from the Standardization of Alzheimer’s Blood Biomarkers (SABB) working group. Alzheimer’s & Dementia. 2022;18(8):1484–1497. doi:10.1002/alz.12510
- Welsh JA, Goberdhan DCI, O’Driscoll L, et al. Minimal information for studies of extracellular vesicles (MISEV2023): From basic to advanced approaches. Journal of Extracellular Vesicles. 2024;13(2):e12404. doi:10.1002/jev2.12404
- Mattsson-Carlgren N, Palmqvist S, Blennow K, Hansson O. Increasing the reproducibility of fluid biomarker studies in neurodegenerative studies. Nature Communications. 2020;11(1):6252. doi:10.1038/s41467-020-19957-6
- Ioannidis JPA, Panagiotou OA. Comparison of effect sizes associated with biomarkers reported in highly cited individual articles and in subsequent meta-analyses. JAMA. 2011;305(21):2200–2210. doi:10.1001/jama.2011.713
- Zhang N, Gu D, Meng M, Gordon ML. TDP-43 is elevated in plasma neuronal-derived exosomes of patients with Alzheimer’s disease. Frontiers in Aging Neuroscience. 2020;12:166. doi:10.3389/fnagi.2020.00166
- Lucien F, Gustafson D, Lenassi M, et al. MIBlood-EV: Minimal information to enhance the quality and reproducibility of blood extracellular vesicle research. Journal of Extracellular Vesicles. 2023;12(12):e12385. doi:10.1002/jev2.12385

Leave a Reply