What a mother eats while breastfeeding may leave a measurable signature in her child’s brain years later, according to a small but provocative pilot study published in npj Science of Food. A team of researchers from Boston Children’s Hospital, Harvard Medical School, and collaborating institutions used machine learning to sift through maternal nutrition data, infant characteristics, and family sociodemographic factors, searching for patterns that could predict how children perform on neurocognitive tests when they reach early school age. The results, while preliminary, point toward a future in which individualized nutritional recommendations during pregnancy and lactation could be grounded in predictive models rather than population averages.
The study set out to address three persistent gaps in the science of nutrition and brain development. First, most existing research examines dietary nutrients in isolation, ignoring the fact that nutrients in a real diet do not act independently but interact with one another in complex ways. Second, the bulk of prior work has focused on preterm infants, even though term-born babies account for more than 90 percent of all newborns, leaving a striking knowledge vacuum about the majority of births. Third, nutrition alone rarely tells the whole story; infant factors, family circumstances, and social determinants of health all shape developmental outcomes and deserve a place in any serious analysis.
To test whether advanced data science could help close these gaps, the researchers turned to an existing birth cohort and extracted individual mother-infant dyad features recorded at delivery and during lactation. The pilot analysis involved just 20 mother-child pairs, a deliberately small sample that the authors frame as a feasibility demonstration rather than a definitive finding. The children’s neurocognitive outcomes were assessed in early school age using the Child Development Inventory, a standardized instrument that captures multiple domains of function, including expressive language and letter scores.
The analytical pipeline combined multivariate machine learning with feature selection, a technique designed to identify which variables carry the most predictive weight when many candidates compete for attention. Rather than forcing every nutrient into the model, the feature selection approach allowed the algorithm to nominate a shortlist of candidate predictors. Among the variables selected were phosphatidylcholine, glutathione, and glycemic load, alongside other nutrients in the maternal diet during lactation. Notably, the model also flagged child demographics and family income as candidate predictors, underscoring the authors’ argument that social determinants of health belong in the same analytical frame as biology.
The choice of phosphatidylcholine is scientifically intriguing. This phospholipid is a key structural component of cell membranes and a source of choline, a nutrient long implicated in brain development and neurotransmitter synthesis. Glutathione, meanwhile, is the body’s principal endogenous antioxidant, raising questions about oxidative stress and its role in the developing brain. Glycemic load captures the blood-sugar impact of the maternal diet as a whole, integrating carbohydrate quality and quantity into a single measure. That these three very different dimensions of nutrition emerged from a data-driven selection process, rather than being imposed by hypothesis, illustrates exactly the kind of discovery-oriented potential the researchers hoped to demonstrate.
Perhaps the most striking result was how unevenly the model’s predictive power was distributed across neurocognitive domains. The association between the selected features and childhood outcomes was strongest for expressive language function and weakest for letter scores, both measured by the Child Development Inventory. The prediction error told the same story from the opposite direction: the model’s error was lowest for expressive language, at just 3 percent, and highest for the letter score, at 72 percent. In other words, the lactation-era features the model selected appear to track expressive language outcomes with remarkable precision in this sample, while their relationship to letter-based skills is far more tenuous.
That asymmetry is worth pausing on. Expressive language, the ability to produce words and communicate ideas, is a domain with well-documented sensitivity to early environmental and nutritional influences. Letter scores, which tap into emerging literacy skills, may depend more heavily on later experiences such as schooling, home literacy practices, and instructional quality, factors that occur well after the lactation window the model draws upon. A 72 percent prediction error for letter scores suggests that the delivery and lactation features simply do not contain enough signal to forecast that domain, which is itself an informative negative result. It hints that different neurocognitive functions may have different windows of nutritional sensitivity, or different degrees of malleability to later social and educational inputs.
The authors are careful to position the work as an impetus for future larger studies rather than a conclusion. With only 20 dyads, the risk of overfitting, in which a model memorizes the quirks of a small sample rather than learning generalizable patterns, is substantial, and the researchers themselves describe the analysis as a pilot designed to generate hypotheses. Still, the feasibility demonstration matters. It shows that multivariate machine learning and feature selection can be applied to nutrition-brain questions in term-born children, that the approach can quantify associations across multiple neurocognitive domains simultaneously, and that it can surface candidate predictors, including nutrients and social factors, that merit testing in larger cohorts.
The funding and conflict-of-interest landscape of the study is also transparent. The parent clinical trial that collected the demographic, nutrition, and brain MRI data was funded by Abbott Inc., and Abbott-affiliated co-authors were involved in that first phase only. The early-school-age neurocognitive outcome data were collected by investigators from Boston Children’s Hospital outside the scope of the original trial and without Abbott funding. Abbott-affiliated authors had no role in the second-phase data collection, though they participated in data interpretation and manuscript editing. The work was supported in part by NIH grant R03 HD107124, a Massachusetts Life Science Center Bits to Byte award, and Abbott Nutrition.
If larger studies confirm and extend these findings, the implications could be far-reaching. Personalized nutrition during pregnancy and lactation remains more aspiration than practice, largely because the field lacks predictive tools linking specific dietary patterns to specific developmental outcomes in individual children. A validated model that could estimate, even probabilistically, how maternal nutrition relates to a child’s language trajectory would give clinicians and families a new kind of evidence base. It could also reshape how nutrition science thinks about the problem altogether, moving from single-nutrient studies toward models that respect the interactions among nutrients and the social context in which they are consumed. For now, the message from this pilot is one of cautious excitement: the data-driven approach works well enough to justify bigger, better-powered studies, and the first hints suggest that what mothers eat while nursing may matter most for the words their children will one day say.
Subject of Research: Machine learning prediction of child neurocognitive outcomes from maternal nutrition and sociodemographic factors
Article Title: Predicting early school-age neurocognitive outcomes from maternal nutrition and sociodemographic factors in a pilot machine learning study
Article References: Bao, R., Wu, D., Vyas, R., Kuchan, M. J., Lasekan, J. B., Leyshon, B., Gopinath, A., Sutton, B. P., Grant, P. E., Morton, S. U., & Ou, Y. (2026). Predicting early school-age neurocognitive outcomes from maternal nutrition and sociodemographic factors in a pilot machine learning study. npj Science of Food. https://doi.org/10.1038/s41538-026-01094-x
Image Credits: AI Generated
DOI: 10.1038/s41538-026-01094-x
Keywords: maternal nutrition, lactation, machine learning, neurocognitive development, expressive language, phosphatidylcholine, glutathione, glycemic load, child development inventory, social determinants of health, feature selection, birth cohort

