“Can stem cells cure ALS?” — anyone who has ever placed even a little hope in regenerative medicine has probably said this to themselves at least once. Amyotrophic lateral sclerosis (ALS) is a disease in which motor neurons break down bit by bit, and there is still no drug that stops its progression. That is precisely why the idea of “taking stem cells out of your own bone marrow, coaching them to pour out nourishment, and returning them to the space around the spinal cord” has been talked about as an intuitively graspable hope — like burying an irrigation pipe directly into a garden that is drying out.
Today’s paper takes that hope and runs it through the coldest testing apparatus medicine possesses: a rigorous phase 3 randomized controlled trial (RCT). The cell therapy is called NurOwn®, formally “mesenchymal stem cells induced to secrete high levels of neurotrophic factors (MSC-NTF).” At six US sites, 189 patients received either cells or placebo three times through the lower back, and the two arms were compared under double-blind conditions.
Let me be honest about the most important thing up front. This trial did not meet its primary efficacy endpoint. You cannot write the headline “stem cells worked in ALS” from this paper. And yet I chose it, because it is exactly inside the failure that the materials for talking honestly about regenerative medicine are packed. Once you separate the promising-looking subgroup signal, the wall of statistical significance, and the movement of the biomarkers, the “true current position” of stem cell therapy comes into view.
👦 Student: Wait, you’re opening by saying it didn’t work? Isn’t this supposed to be a hopeful story?
🧬 Dr. Exotaro: Hope and fact are different things. Until you put the facts down accurately, you can’t even see where the real hope lies. Today we won’t stop at “it failed” — we’ll read all the way through to “so how should the next one be designed?”
Journal Information
- Paper title: A randomized placebo-controlled phase 3 study of mesenchymal stem cells induced to secrete high levels of neurotrophic factors in amyotrophic lateral sclerosis
- Authors: Merit E. Cudkowicz, Stacy R. Lindborg, Namita A. Goyal, Robert G. Miller, Matthew J. Burford, James D. Berry, Katharine A. Nicholson, Tahseen Mozaffar, Jonathan S. Katz, Liberty J. Jenkins, Robert H. Baloh, Richard A. Lewis, Nathan P. Staff, Margaret A. Owegi, Donald A. Berry, Yael Gothelf, Yossef S. Levy, Revital Aricha, Ralph Z. Kern, Robert H. Brown Jr, Anthony J. Windebank
- Affiliations: The Healey Center at Massachusetts General Hospital / Harvard Medical School (Boston, MA), the University of California Irvine, Cedars-Sinai Medical Center, Mayo Clinic, University of Massachusetts Medical School, Berry Consultants, and the trial sponsor Brainstorm Cell Therapeutics (New York, NY / Petah Tikva, Israel), among others
- Journal: Muscle & Nerve (published by Wiley Periodicals LLC), Volume 65, Issue 3, pages 291–302, 2022
- DOI / link: 10.1002/mus.27472 (https://doi.org/10.1002/mus.27472) / clinical trial registration NCT03280056 (the BCT-002 trial)
- Impact Factor: roughly around 3 (approximate; citation needed. The Clarivate JCR value varies from year to year, and it is treated here as an outside estimate not stated in the original paper). A specialty journal covering neuromuscular disease, peripheral nerve, and electromyography
- Open access: Creative Commons Attribution-NonCommercial license (CC BY-NC). © 2021 BrainStorm Cell Therapeutics. Muscle & Nerve published by Wiley Periodicals LLC. Use, distribution, and reproduction are permitted for non-commercial purposes only, provided the work is properly cited
- Review timeline: Received October 8, 2021 / Revised December 4, 2021 / Accepted December 7, 2021
- Ethics approval: Informed consent was obtained and documented from all participants. Institutional Review Board (IRB) approval was obtained at all sites. The authors confirm compliance with the journal’s ethical publication policy
- Funding: Brainstorm Cell Therapeutics funded the trial. Additional support came from the California Institute for Regenerative Medicine (CIRM, CLIN2-09894), the ALS Association, and I AM ALS. Funders other than Brainstorm had no role in data collection, analysis, interpretation, or manuscript preparation
- Conflicts of interest: Substantial conflicts of interest exist (a sponsor-led trial). S.R.L., R.K., Y.G., Y.S.L., and R.A. are Brainstorm employees and hold Brainstorm stock. D.A.B. received personal fees from Brainstorm. Corresponding author M.C. and many other authors received honoraria, research support, or consulting fees from numerous companies including Biogen, MT Pharma, and Denali. N.A.G. received research support from Brainstorm. On the other hand, A.J.W., L.J., M.A.O., M.B., R.M., and T.M. reported no conflicts of interest to disclose
- Data availability: De-identified patient data are not publicly shared (because they are the subject of ongoing regulatory interactions toward possible future approval)
About the corresponding author: The corresponding author is Merit E. Cudkowicz (MD, MSc). She is based at Massachusetts General Hospital, serves as the Julieanne Dorn Professor of Neurology at Harvard Medical School, and directs the hospital’s Sean M. Healey & AMG Center for ALS (contact email cudkowicz.merit@mgh.harvard.edu). She is one of the world’s leading figures in ALS clinical research, known for co-founding the Northeast ALS Consortium (NEALS) and for leading the HEALEY ALS Platform Trial, which evaluates multiple drugs simultaneously. It is worth remembering that the clinical scientist who has most systematically kept challenging “a disease with no treatment” is reporting the result of this trial head-on — even though it is negative.
What Was Not Known Until Now?
ALS is a neurodegenerative disease in which the motor cortex of the brain and the motor neurons of the spinal cord progressively degenerate and are lost. Strength drains from the limbs, and eventually the ability to speak, swallow, and even breathe is taken away. There is no treatment that stops or reverses the progression, and it remains one of the largest “unmet needs” in medicine.
Why is it so difficult? One reason is that the pathology of ALS is not “a single broken gear” but a complex system in which multiple failures advance simultaneously. On top of the degeneration of the nerve cells themselves, neuroinflammation is now thought to be heavily involved. Immune runaway, depletion of trophic factors, abnormal protein aggregation — many pathways interlock to corner the motor neurons. If that is the case, wouldn’t a “multi-pronged therapy that acts on several pathways at once” make more sense than a conventional drug aimed at a single molecule? This is where cell therapy enters the stage.
👦 Student: Why make “the cells themselves” into the drug? Why won’t an ordinary medicine do?
🧬 Dr. Exotaro: If an ordinary drug is “a courier delivering a single ingredient,” a mesenchymal stem cell is closer to “a resident staff member who moves in on site and keeps releasing many kinds of substances while reading the situation.” It secretes trophic factors, it calms inflammation — it tries to fix the environment with many moves rather than one. In a disease like ALS where several pathways are broken, that “many moves” quality is what people pin their hopes on.
MSC-NTF (NurOwn), the protagonist of this paper, pushes that idea one step further. Autologous (your own) mesenchymal stem cells (MSCs) harvested from adult bone marrow are induced, under a proprietary ex vivo culture condition, “to secrete neurotrophic factors (NTFs) at high levels” — an approach that puts the resident staff through special training in advance, maximizing their “ability to release nourishment,” before returning them to the body.
Earlier-stage studies — the phase 1/2 and phase 2 trials (Petrou et al. 2016, Berry et al. 2019) — had suggested safety, preliminary efficacy, and favorable changes in CSF (cerebrospinal fluid) biomarkers after a single intrathecal administration. But these were small, weakly controlled, exploratory stages. A “looks good” from a small trial has about the precision of a weather forecast saying “it might rain.” Placebo effects, chance, and bias in patient selection often make a difference appear where none exists. That is exactly why we use the apparatus of “large numbers, double blinding, placebo control” to sort chance from the real thing. “Can repeated intrathecal MSC-NTF safely slow ALS progression under the most rigorous conditions of a phase 3 trial?” — nobody yet had the answer to this core question. The BCT-002 trial we look at today was assembled to settle it.
What Did This Paper Find?
First, let’s pin down the shape of the trial. BCT-002 was a randomized, double-blind, placebo-controlled, parallel-group phase 3 trial. Eligible patients had ALS meeting the revised El Escorial criteria, a total score on the revised ALS Functional Rating Scale (ALSFRS-R) of 25 or higher at screening, symptom onset within 24 months before screening, an upright slow vital capacity (SVC) of at least 65% of predicted, and an age of 18–60. After an 18-week run-in period (which included an outpatient bone marrow aspiration to manufacture the autologous cells), a final condition before randomization required “a decline of at least 3 points on the ALSFRS-R.” The design enrolled only patients who were demonstrably progressing.
Eligible participants were allocated 1:1 to the MSC-NTF arm or the placebo arm and received three intrathecal administrations by standard lumbar puncture at weeks 0, 8, and 16, followed by 12 weeks of observation (CSF was collected seven times from each participant, and cell manufacturing was handled by the Dana-Farber Cancer Institute or City of Hope). Between August 2017 and September 2020, 263 people were screened, 196 were allocated, and 189 received at least one treatment (of whom 7 were untreated and 45 discontinued early). The mean age of participants was 49, 67% were male, and the mean baseline ALSFRS-R was 31.
There is one piece of foreshadowing here that shapes the later interpretation. Compared with other late-phase ALS trials, this trial included a larger share of more advanced (more severely affected) patients, and some patients already scored 0 on individual items (meaning that function was already lost and further progression could not be shown).
Primary Endpoint: Not Met
The primary efficacy endpoint was a “responder analysis.” The rate of change per month in ALSFRS-R was derived by linear regression, with separate regression lines fitted to the pre-treatment period and to the post-treatment period through week 28, and a patient whose post-treatment progression improved by at least 1.25 points/month relative to the pre-treatment period was defined as a “responder” (clinical response) (participants who died of disease progression were counted as non-responders).
The result is unambiguous. At week 28, the responder criterion was met by 33% (31/95) in the MSC-NTF arm and 28% (26/94) in the placebo arm. The odds ratio (OR) was 1.33 (95% confidence interval 0.63–2.80), P=.45. There was no statistically significant difference, and the primary endpoint was not met. A prespecified sensitivity analysis counting all deaths as non-responders gave an identical result. Because the authors used a hierarchical testing plan (testing in sequence to prevent false positives from multiple comparisons), once this primary test failed, every subsequent P value is nominal and does not control error at the .05 level — do not let go of this caveat all the way to the end.
The key secondary endpoint (improvement of the slope by 100% or more, i.e., the proportion in whom progression essentially stopped) told the same story. MSC-NTF 13.7% (13/95), placebo 13.8% (13/94), OR=0.998 (95%CI 0.42–2.40), P=0.997. Both arms were around 14%, with no difference. Let’s look at the other secondary endpoints as well.
- Least squares (LS) mean change in total ALSFRS-R score from baseline to week 28: MSC-NTF 5.52 (standard error 0.67) vs placebo 5.88 (0.67), LS mean difference 0.37 (95%CI −1.47, 2.20), P=.69
- Combined Analysis of Function and Survival (CAFS), week 28 LS mean: MSC-NTF 73.74 vs placebo 72.21, LS mean difference 1.53 (95%CI −10.65, 13.72), P=.80
- LS mean change from baseline in SVC (% predicted): MSC-NTF 12.94 vs placebo 11.55, P=.56
- Event-free probability for death from disease progression (Kaplan-Meier, through week 32): 90.43% vs 92.24%, P=.21
- Event-free probability for all-cause mortality (same): 88.32% vs 89.17%, P=.11
None of the secondary endpoints showed a statistically significant between-group difference. Differences in death and tracheostomy were not significant either, and event-free probabilities all exceeded 88%. These are the skeletal facts of this trial.
👦 Student: 33% versus 28% — that looks like a 5% difference to me… and you still say there’s no difference?
🧬 Dr. Exotaro: This is the crux of statistics. Don’t be fooled by the look of “5%.” The 95% confidence interval of the odds ratio runs “from 0.63 to 2.80” — that is, it includes “possibly worse than placebo (0.63)” all the way to “possibly 2.8 times better.” Straddling 1.0 (= no difference) means “even if the true difference were zero, a gap of this size could happen by chance.” So you cannot puff out your chest and say “it worked.”
Prespecified Subgroup: A Signal, but Not Significant
Here is where the “hope” and the “caution” of this paper cross paths. At the planning stage, the authors had prespecified a subgroup split at “baseline ALSFRS-R of 35 (= the assumed mean value).” The hypothesis was that the effect might be easier to see in less-impaired patients.
In that less-impaired subgroup (baseline ALSFRS-R ≥35, n=58), a responder rate of 35% (9/26) in the MSC-NTF arm vs 16% (5/32) in the placebo arm was observed. On the face of it, that is more than a two-fold difference. However — OR=2.6 (95%CI 0.45–14.36), P=.29. It is not statistically significant. The extreme width of the confidence interval, “0.45 to 14.36,” reflects the fact that only 31% of all participants fell into this group, so the sample is too small and the estimate becomes unstable (there is not enough power). In the more severely affected subgroup (baseline <35, MSC-NTF n=69, placebo n=62), by contrast, MSC-NTF 31.9% (22/69) vs placebo 33.9% (21/62), OR=0.87, P=.74 — the response rates were essentially the same.
Furthermore, if the boundary is redrawn at the actually observed mean (ALSFRS-R ≥31), the figures become MSC-NTF 35.4% vs placebo 15.4%, which is “nominally significant (P<.05).” But this is a nominal P value obtained after the hierarchical testing had already broken down, from an analysis that used the observed value after the fact rather than the pre-assumed 35, and error is not controlled at α=.05.
👦 Student: In the less-impaired patients it’s 35% vs 16%. Surely that at least is evidence that it’s working?
🧬 Dr. Exotaro: I understand the feeling. But this is not “evidence” — it is “the seed of a hypothesis” (hypothesis-generating). If you flip a coin a few times and heads happens to come up several times in a row, you can’t conclude “this coin favors heads,” right? A confidence interval that stretches out to 14-fold is precisely statistics screaming “we still can’t say anything.” These numbers are a map telling us where to dig next; they are not the treasure itself.
Biomarkers: The Cells Really Were Doing Something
This is why I don’t want to file this paper away as “just a failure.” There was no difference in the clinical scores (patient function), but the CSF biomarkers (indicators of the biological changes happening inside the body) showed large, statistically significant changes. And they did not move in the placebo arm.
- Vascular endothelial growth factor (VEGF): a neurotrophic factor that protects motor neurons. It doubled from baseline in the MSC-NTF arm at week 20 and was significantly higher than placebo at every time point (P<.05). At week 20 it was 258% of placebo; overall treatment effect P<.0001
- Monocyte chemoattractant protein-1 (MCP-1): a chemokine involved in neuroinflammation that correlates with ALS progression. It was significantly reduced by MSC-NTF at every time point (P<.05), overall treatment effect P<.0001. At week 20 it was 74% of placebo
- Neurofilament light chain (NfL): a degeneration marker reflecting axonal damage. At week 20, MSC-NTF was 82% of placebo
VEGF goes up (more trophic support), MCP-1 goes down (inflammation calms), NfL goes down (the nerves break down more slowly) — every direction is exactly what the treatment intended. In pharmacological terms this is “target engagement”: evidence that the administered cells really did reach the biological pathways they were aimed at. Moreover, a model combining these biomarkers with MSR1, Fetuin-A, and the ENCALS (European Network to Cure ALS) risk score predicted clinical response with 82.5% accuracy (based on the ROC curve).
Safety matters too. Repeated intrathecal MSC-NTF was generally well tolerated, and there were zero deaths related to the study treatment. There were 16 deaths after randomization (14 post-treatment, 2 pre-treatment), but none were judged related to the study treatment, and they were concentrated among advanced patients. Treatment-emergent adverse events (TEAEs) such as pain associated with lumbar puncture or bone marrow aspiration were somewhat more frequent in the MSC-NTF arm (procedure-related TEAEs: MSC-NTF 93.7% vs placebo 87.2%), but it is reasonable to read this as a reflection of “the burden of the administration procedure” rather than “toxicity of the cells.”
👦 Student: The biomarkers moved, but there was no difference in the patients’ function. Isn’t that a contradiction?
🧬 Dr. Exotaro: It’s not a contradiction — it’s a distance. The biomarkers are, so to speak, “the sound of the engine actually starting.” Whether the car actually drove to the finish line (whether the patient’s function was preserved) is another matter, and if the road is too rough, it won’t move. This time the roughness of the road — “many patients had progressed too far, and the rating scale was pinned near its floor” — may have prevented a genuine biological effect from being picked up as a clinical difference. That is exactly why, next time, the key is to use “the sound of the engine” as a guide and choose a road that can be driven.
How Will the Future Change? (The Road to the Clinic)
The conclusion the authors drew from this trial was restrained. “Primary and secondary endpoints did not reach significance in the overall population. However, a prespecified subgroup suggested that less-impaired MSC-NTF patients may have better preserved function, and biomarkers provided evidence of target engagement consistent with the mechanism of action. Given the unmet patient need, these results warrant further investigation” — neither “cured” nor “meaningless,” but a stance that builds a bridge to the next step out of an instructive failure.
The concrete “lessons for next time” fall into three broad categories. First, choose a more homogeneous, less severely affected patient population. Including many patients who had progressed too far is thought to have produced a “dilution effect” that watered down the overall effect, and the authors suggest using ALSFRS-R >25 at baseline, rather than at screening, as the enrollment criterion. Second, account for the “floor effect” of the ALSFRS-R (in severely affected patients near the lower limit, slowed progression can be misclassified as “improvement”). Third, use biomarkers to predict treatment response. The idea is that a model with 82.5% predictive accuracy could become a tool for selecting “patients likely to respond” in advance in future trials (a detailed analysis of the correlation between CSF biomarkers and the primary outcome is to be published separately).
That said — it would be dishonest not to also record what actually happened outside this paper. What follows is not contained in this paper (published in 2022); it is subsequent regulatory history, and should be read as external information based on primary sources such as the public materials of the FDA advisory committee. After this trial, the sponsor BrainStorm Cell Therapeutics submitted a Biologics License Application (BLA) for NurOwn to the FDA, but on September 27, 2023, the FDA’s Cellular, Tissue, and Gene Therapies Advisory Committee (CTGTAC) voted 1 in favor, 17 against, and 1 abstention that it did not provide substantial evidence of effectiveness in mild to moderate ALS. In response, BrainStorm announced its intent to withdraw on October 18, 2023, and withdrew the BLA on November 3 of the same year (the FDA characterized it as a withdrawal “without prejudice,” leaving room for a future resubmission). None of this falls within the verification scope of the original paper (2022), and for details it is safest to consult the public materials of that advisory committee and the company’s timely disclosures.
👦 Student: So in the end, NurOwn couldn’t become an ALS treatment after all…
🧬 Dr. Exotaro: “It has not been approved at this point in time” is the accurate way to put it. The story is neither “failed and over” nor “about to be approved”; it sits at that most frustrating midpoint, “promising but unestablished.” And that is exactly why our job as medical professionals is neither to fan expectations nor to throw cold water, but to point accurately at where we are right now.
Dazzling numbers from preclinical work and small trials being brought to a halt, for the moment, at the gate of a large phase 3 trial and regulatory review — that is not unusual; if anything, it is evidence that rigorous verification is working. How much of these lessons the next trial can implement is the question. Whether cell therapy can establish clinical value in ALS depends on that design.
How to Read This Study Critically (Limitations, and the Road to Better Quality)
Let’s begin fairly. There is an honesty in this paper that deserves praise. It does not hide the negative primary result; it states plainly in the abstract at the very top that “the primary endpoint was not met.” For a sponsor-led trial, that is a mark of scientific conscience. The clear separation of prespecified subgroups from post hoc analyses, and the repeated warnings that the P values are nominal, likewise show high statistical literacy.
On that basis, let’s look at the limitations.
First, the biggest problem is that patients who had progressed too far were enrolled. The mean baseline ALSFRS-R was 31, and some patients already scored 0 on individual items. As the authors themselves acknowledge, this markedly reduced the ability to detect a treatment effect in the overall population. The “floor effect” of the ALSFRS-R — in patients pinned to the lower limit of the scale, the primary endpoint’s definition of slowed progression does not function properly — makes the results harder to interpret. A threshold-based primary endpoint (improvement of 1.25 points/month) also has the mathematical quirk that “the faster a patient progresses, the easier it is to meet it,” and in a more homogeneous population it would have behaved more consistently.
Second, the subgroup signal is nothing more than an observation made under insufficient power. Only 31% of participants fell into the prespecified ALSFRS-R ≥35 subgroup, so power dropped sharply. The authors themselves state honestly that “even if baseline ALSFRS-R >25 had been used as the enrollment criterion, a trial of the same size (n=196) would likely have been underpowered (estimated power 60%).” In short, there simply were not enough participants, from the design stage onward, to test the effect in a less-impaired subgroup.
Third, the weight of the conflicts of interest. This was a trial funded and led by BrainStorm, and several authors are employees and shareholders of the sponsor. That in itself does not imply fabrication of results, and an independent data safety monitoring board was in place. Still, one should be cautious about taking the authors’ forward-leaning phrasing in the text — expressions such as “suggest a potential treatment effect” — at face value as a conclusion. The statement that post hoc imputation of missing data made “the treatment difference more than 45% larger” is likewise, being an exploratory analysis, best treated as a hypothesis rather than a confirmation.
Fourth, the impact of COVID-19. A protocol amendment in March 2020 allowed remote visits, which had a major effect particularly on the collection of SVC assessments. This is a constraint shared by trials run during the pandemic, and it may have compromised data completeness.
So, how could this have been an even higher-quality study? The answer overlaps almost exactly with what the paper itself offers as “lessons for next time.” (1) Set the enrollment criterion at “baseline ALSFRS-R >25 (or even less impaired)” to avoid the floor effect and assemble a homogeneous population. (2) Design from the outset a sample size large enough to secure the necessary power in that population. (3) Use a continuous-measure primary endpoint that is less vulnerable to floor effects. (4) Use biomarker-based patient selection (enrichment) to narrow down to “patients likely to respond.” (5) Where possible, run the primary analysis in parallel by an analysis team independent of the sponsor. A trial that satisfies these is the one that could finally give a clear answer to the question “do stem cells work in ALS?”
👦 Student: Listening to all these criticisms, I’m starting to feel like the trial itself was no good.
🧬 Dr. Exotaro: Please don’t misunderstand that. Pointing out weaknesses in a design is not the same as denying the value of a trial. This trial was, if anything, a courageous undertaking in that it rigorously tested a cell therapy in a real-world patient population that included advanced cases. The lessons obtained — which patients, on which scale, in what numbers should be measured — are worth their weight in gold coins for the next trial. Science climbs by using these “meaningful failures” as stairs.
Dr. Exotaro’s Perspective
To be honest, this paper is not someone else’s business for me. I too have worked on regenerative medicine for spinal cord injury (SCI) and stroke using mesenchymal stem cells and the extracellular vesicles (EVs / exosomes) they secrete. The feeling behind the metaphor of MSCs as “resident staff who fix the environment with many moves” is exactly what I experience day to day, and I receive the result of this phase 3 trial with both joy and caution.
Start with the joy. The fact that the CSF biomarkers moved clearly, and separated distinctly from placebo, is important evidence supporting the biological rationale of cell therapy. That movement matches precisely the scenario we have observed in animal models and human preclinical work — “MSCs/EVs calm inflammation, supply nourishment, and slow the breakdown of nerves.” The cells really were doing their job inside a living body: that is a finding I find straightforwardly encouraging.
At the same time, the caution. Between “it is working biologically” and “a patient’s life changes,” there is still a deep valley. Even when preclinical data look beautiful, there is no guarantee whatsoever that they translate into functional improvement in humans. Treating a non-significant signal (P>0.05) as “evidence that it worked” risks becoming a betrayal of patients. The subgroup numbers, 35% vs 16%, certainly make my heart beat faster. But the moment I speak about them with the caveat OR=2.6, P=.29 stripped away, I stop being a physician and become a salesman.
The world of regenerative medicine has a structural problem in which expectations tend to run ahead. The words “stem cells” and “exosomes” have a power to attract people all by themselves. That is exactly why a stance of calmly separating success, failure, and lessons learned is decisively necessary. In spinal cord injury and stroke, we stand before the same gate. The cells and EVs are doing something — that much I can believe. The question is whether that “something” can be measured in the right patients, with the right indicators, in sufficient numbers, and shown to clear the statistical wall. Do not abandon hope, but do not exaggerate — I believe that narrow path is the only way to make regenerative medicine real.
👦 Student: So your own research has the same “valley,” doesn’t it.
🧬 Dr. Exotaro: It certainly does. And that is why I sincerely respect these authors for publishing their negative result head-on. Not hiding inconvenient results — that is the shortest route to eventually arriving at a treatment that truly works. Just as you today did not swallow “the difference between 33% and 28%” whole, I hope you keep the eye that reads past the number to the confidence interval behind it. That is exactly the eye that protects patients.
