In the race to feed a warming planet, scientists have long sought to decode the genetic instructions that make some crops flourish while others falter. Now, an international team of researchers has taken a major step forward in that quest, mapping the intricate genetic architecture that governs fruit weight in tomato—one of the world’s most economically important vegetable crops. The study, published in the journal Theoretical and Applied Genetics, reveals that the size of the tomato on your kitchen counter is overwhelmingly dictated by its DNA rather than by the environment in which it grew, and that the evolutionary journey from wild berry to beefsteak has left detectable fingerprints throughout the tomato genome.
The research, led by scientists affiliated with multiple institutions across the United States and Europe, harnessed the power of the Varitome population—a curated collection of 166 tomato accessions that spans the full arc of tomato domestication. This panel includes 28 accessions of the wild progenitor Solanum pimpinellifolium (SP), 117 accessions of the semi-domesticated Solanum lycopersicum var. cerasiforme (SLC), and 21 accessions of the fully cultivated Solanum lycopersicum var. lycopersicum (SLL). By studying these three groups side by side, the team was able to trace how thousands of years of human selection have sculpted the genetic landscape of a crop that now ranks among the most valuable horticultural commodities on Earth.
The team grew all accessions across four dramatically different environments: Antalya in Türkiye, Valencia in Spain, and Florida and Georgia in the United States. These sites represent contrasting agroecological zones, providing a rigorous test of how stable the genetic control of fruit weight truly is. The plants were arranged in randomized complete block designs, and fruit weight data were collected and normalized using a base-10 logarithmic transformation to stabilize variance. The researchers then applied a multi-environment linear mixed model, implemented through restricted maximum likelihood, to partition the observed phenotypic variance into genetic, environmental, genotype-by-environment interaction, and residual components.
What they found was striking. Genotype alone explained more than 92 percent of the phenotypic variance in fruit weight, while environment and genotype-by-environment interaction contributed only marginal effects. The broad-sense heritability estimate reached 0.98, and the overall model fit was exceptionally high with an R-squared of 0.96. Within individual replicated environments, heritability estimates exceeded 0.97 in both Florida and Georgia. These numbers confirm that tomato fruit weight is governed primarily by stable additive genetic effects rather than by environmental plasticity—a finding that has immediate practical implications for breeders attempting to predict fruit weight across different growing regions and seasons.
But the researchers did not stop at confirming what many suspected. Their next move was to dissect the molecular machinery underlying this strong genetic control, and here they deployed an approach that goes well beyond the standard toolbox of modern plant genetics. Rather than relying solely on single nucleotide polymorphisms—the workhorse markers of most genome-wide association studies—the team analyzed three distinct classes of genomic variation simultaneously: SNPs, insertions and deletions (INDELs), and structural variants (SVs). SNPs represent single-letter changes in the DNA sequence, while INDELs involve the insertion or deletion of one to several hundred base pairs. Structural variants are even larger genomic rearrangements—deletions, duplications, inversions, and transpositions that can span thousands of base pairs and alter gene dosage, chromatin structure, or regulatory sequences.
This multi-variant strategy proved powerful. Through genome-wide association analyses conducted across all four environments using both the TASSEL and BLINK statistical frameworks with false discovery rate correction, the team identified 15 significant SNP loci, 10 INDEL loci, and 10 SV loci associated with fruit weight. Crucially, several of these associations co-localized with genes already known to govern tomato fruit size—including fw2.2/CELL NUMBER REGULATOR, fw3.2/KLUH, fw11.3/CELL SIZE REGULATOR, lc/WUSCHEL, and fas/CLAVATA3. These genes operate through diverse biological mechanisms, from regulating cell division during fruit development to controlling the organization of the floral meristem, the structure that determines how many seed-producing compartments a fruit will contain.
Among all the associations detected, chromosome 5 stood out as particularly notable. A 390-base-pair deletion at position 45,551,023 was detected in all four environments, while an INDEL at position 55,361,233 appeared in three of the four environments. The consistency of these signals across geographically and climatically distinct trial sites suggests they represent stable, improvement-associated candidate loci that could be validated for use in marker-assisted selection programs. For breeders, identifying such robust markers is invaluable because it means the markers can be trusted regardless of where or when the crop is grown.
The study also revealed a fascinating evolutionary pattern. When the researchers tracked allele frequencies across the three domestication groups—from wild SP through semi-domesticated SLC to fully cultivated SLL—they observed clear and consistent shifts at multiple loci. Some alleles showed progressive frequency changes consistent with selection during domestication and subsequent crop improvement, while others were detected almost exclusively in cultivated germplasm, suggesting more recent origins tied to modern breeding programs. This trajectory is consistent with prior genomic work showing that domesticated SLL tomatoes retain only about 22.6 percent of the standing genetic diversity found in wild SP populations, while semi-domesticated SLC accessions retain approximately 53.8 percent. Severe bottlenecks during domestication and targeted selection for larger fruits have narrowed allelic diversity precisely in the genomic regions that control fruit weight, creating a paradox for breeders: the genes that matter most are also those with the least remaining variation in elite germplasm.
To translate these genetic insights into practical breeding tools, the team built five distinct genomic prediction models. Model 1 used only SNPs, Model 2 only INDELs, Model 3 only SVs, Model 4 combined all three marker types additively, and Model 5 incorporated genotype-by-environment interaction kernels on top of the combined markers. Each model was tested using four complementary cross-validation scenarios that simulate different breeding situations: CV1 predicts tested genotypes in tested environments, CV2 predicts untested genotypes in tested environments, CV0 predicts tested genotypes in untested environments, and CV00—the most demanding scenario—predicts untested genotypes in untested environments. Each scenario was repeated 100 times with random 70-30 training-testing splits to ensure robust estimates.
The prediction algorithms themselves were also diverse, spanning four analytical frameworks: Bayesian Genomic Linear Regression (BGLR), Partial Least Squares regression (PLS), Random Forest (RF), and Deep Learning (DL) neural networks. BGLR used reproducing kernel Hilbert space regression with 5,000 MCMC iterations, PLS employed kernel-based latent variable decomposition through the SKM package, Random Forest used 500 regression trees, and the deep learning model used a fully connected neural network with ReLU activation, L2 regularization, dropout, and the Adam optimizer implemented through TensorFlow and Keras.
When the dust settled from thousands of model runs, two patterns emerged clearly. First, the choice of validation scenario and prediction algorithm mattered far more than the type of genetic marker used. Partial Least Squares and BGLR consistently outperformed Random Forest and Deep Learning models, particularly under the most stringent cross-validation scenarios involving untested genotypes and untested environments. Second, while differences among marker classes were generally modest, INDEL-based models frequently achieved the highest prediction accuracies under the most challenging scenarios. This is a noteworthy finding because INDELs, especially those falling within coding sequences or promoter regions, are more likely than SNPs to have direct functional consequences on gene expression or protein structure. Their predictive power suggests they may be capturing signals that SNPs alone cannot fully resolve.
The integration of all three variant types into a single model—Model 4—improved biological interpretation and enabled more comprehensive identification of domestication- and improvement-associated loci than any single marker class could achieve alone. The researchers also estimated genomic heritability separately for each marker class using Bayesian models with marker-specific genomic relationship matrices, finding that each class captured a meaningful but partially overlapping slice of the genetic variance underlying fruit weight.
Principal component analysis of the SNP data painted a vivid picture of the domestication gradient. The first principal component, explaining 20.12 percent of total genetic variation, cleanly separated wild SP accessions from both SLC and SLL groups. The second component, accounting for an additional 8.65 percent, captured diversity within species, particularly among the genetically diverse wild accessions. SLL accessions clustered tightly near the center of the plot, reflecting the severe reduction in genetic diversity that accompanied domestication and modern breeding, while SLC occupied an intermediate position, consistent with its role as a transitional population.
Cross-environment correlation analyses added another layer of biological insight. Fruit weight showed strong correlations across environments for SLC and SLL accessions, with Pearson correlation coefficients ranging from 0.76 to 0.99, indicating that genotype rankings remain stable across different growing conditions. In contrast, wild SP accessions exhibited low or non-significant correlations ranging from 0.37 to 0.72, revealing that their fruit weight is highly environment-dependent. Finlay-Wilkinson regression confirmed this pattern: cultivated SLL accessions displayed nearly parallel reaction norms with shallow slopes, indicating strong environmental buffering, while wild SP accessions showed steep and highly variable slopes, demonstrating pronounced environmental sensitivity.
The implications of this work extend well beyond the tomato field. By demonstrating that a multi-variant genomic framework—incorporating SNPs, INDELs, and SVs alongside genotype-by-environment interaction kernels—can simultaneously dissect genetic architecture and deliver accurate predictions across environments and genetic backgrounds, the study provides a template that could be applied to other crops shaped by similar domestication histories. For a world that will need to produce substantially more food on limited land under increasingly unpredictable climatic conditions, the ability to predict complex traits like fruit weight with precision, using genomic data rather than laborious field phenotyping, represents a meaningful acceleration of the breeding cycle. The humble tomato, it turns out, has much to teach us about the future of agriculture.
Subject of Research: Genetic architecture and genomic prediction of tomato fruit weight across the domestication continuum using multi-variant GWAS (SNPs, INDELs, and structural variants)
Subject of Research: Agriculture
Article Title: Multi-variant GWAS and genomic prediction dissect the genetic architecture underlying tomato fruit weight
Article References: Topcu, Y., Adak, A., Kayikci, H. C., Aydin, S., Yildiz, K., Ramos, A., Tieman, D. M., Visa, S., van der Knaap, E., & Sapkota, M. (2026). Multi-variant GWAS and genomic prediction dissect the genetic architecture underlying tomato fruit weight. Theoretical and Applied Genetics, 139(9), Article 259. https://doi.org/10.1007/s00122-026-05326-2
Image Credits: AI Generated
DOI: 10.1007/s00122-026-05326-2
Keywords: tomato, fruit weight, GWAS, genomic prediction, SNPs, INDELs, structural variants, domestication, Solanum pimpinellifolium, Solanum lycopersicum, genotype-by-environment interaction, heritability, Varitome
Cite Scienmag News
APA MLA Chicago
Juliet Wilcox. (September 7, 2026). GWAS and genomic prediction reveal genes controlling tomato fruit weight. Scienmag. https://scienmag.com/gwas-and-genomic-prediction-reveal-genes-controlling-tomato-fruit-weight/
Copy citation Download RIS
Tags: Advances in tomato quantitative trait analysisCurated tomato accessions for genetic studyEvolutionary fingerprints in tomato genomeGenetic architecture of tomato sizeGenetic architecture of tomato yield traitsGenetic markers for tomato breedingGenome-wide association studies in tomatoesgenomic prediction in crop breedingGenomic prediction of crop traitsIdentification of fruit size genes in tomatoIdentification of genes controlling tomato sizeImpact of domestication on tomato genomeInfluence of DNA on tomato fruit sizeRole of DNA in tomato size determinationRole of GWAS in crop improvementTomato accessions and genetic diversityTomato breeding for fruit weightTomato domestication and evolutionTomato fruit weight geneticsWild and cultivated tomato genetic diversityWild versus cultivated tomato genetics

