Strong & Fast

Methodology

A percentile is meaningless until you name the population — and for most tests of athleticism, nobody has ever measured the population. This page explains exactly what we did about that, test by test, and how much to trust each number.

Evidence policy · Dated changes and their scope

What the categories measure

The 30 m sprint measures straight-line acceleration. We group it with jumps under Power to keep the calculator to four categories; sprint speed and jump power remain distinct abilities. Agility covers planned change of direction with the 5-10-5 shuttle and agility T-test, not reactions to an unpredictable stimulus. Supported sprint estimates count toward Power; the four categories retain equal weight.

The two comparison groups

General population

All adults of your sex and age — not gym-goers, not app users, not "people who train". Most adults have never attempted a one-rep max and a large share cannot jog a continuous mile, so anyone who trains consistently climbs this scale quickly. That is the honest answer, not a flaw: against everyone, training counts for a lot.

NCAA Division I athletes

A single pooled group per sex. Where sport-level comparisons are supported, each sport enters the pool weighted by its actual head-count from the NCAA's 2024-25 participation report (92,285 men and 71,589 women after removing athletes double-counted across indoor track, outdoor track and cross country). The largest shares are, for men: football 38%, baseball 14%, track 13%, soccer 7%, basketball 6%, swimming 4%; for women: track 20%, soccer 15%, softball 10%, volleyball 9%, swimming 8%, rowing 8%. Football and track are split into position and event groups, because a lineman and a cornerback — or a thrower and a 10k runner — are different populations.

Sports with no published data are not dropped — that would quietly turn "D1 athletes" into "D1 football and basketball players". They are represented by one explicitly modeled component carrying the remaining head-count. The D1 comparison is raw: no age or size adjustment. Weighted pull-ups use one explicitly modeled all-athlete distribution because the evidence cannot yet support a sport breakdown.

Three grades of evidence

Every percentile in the calculator carries one of these markers.

An all-D1 comparison is labeled modeled whenever any part of its athlete pool uses an assumed distribution. A sport with direct data can have a stronger evidence label than the full pool. These labels describe how estimates are built; they are not validated accuracy ratings.

TestDomainGeneral populationD1 menD1 women
DeadliftStrengthModeledModeledModeled
Back squatStrengthModeledModeledModeled
Bench pressStrengthModeledModeledModeled
Overhead pressStrengthModeledModeledModeled
Weighted pull-upStrengthModeledModeledModeled
Barbell row (Pendlay)StrengthModeledModeledModeled
Grip strengthStrengthMeasuredModeledModeled
Vertical jumpPowerDerivedModeledModeled
Standing broad jumpPowerModeledModeledModeled
Single-leg triple hopPowerModeledModeledModeled
5-10-5 shuttleAgilityModeledModeledModeled
Agility T-testAgilityModeledModeledModeled
One-mile runEnduranceDerivedModeledModeled
1.5-mile runEnduranceDerivedModeledModeled
30 m sprintPowerModeledModeledModeled
400 m runEnduranceModeledNo dataNo data

Modeled running and strength estimates

The 30 m sprint adds acceleration and the 400 m run adds speed endurance. Both accept a time alone. Estimates contribute to scores and training priorities where supported. Keep your timing method consistent when retesting.

Sprint population estimates cover ages 18–39, using active-adult studies with explicit assumptions for the wider population and aging into the thirties. D1 sprint estimates combine college football, soccer and lacrosse data with a national-team proxy for other sports. The sport transfers and timing allowances are modeled, not measured NCAA-wide distributions.

The 400 m population estimate converts age- and sex-specific law-enforcement 300 m tables using a paired 300/400 m study. It covers men ages 18–59 and women ages 18–49. The source is a proxy for adults; the conversion for women and slower runners is an assumption. Older-age running comparisons and pooled D1 400 m remain unavailable because the transfer evidence is too weak. Selected NCAA 400 m race lists are shown as specialist references, not all-athlete percentiles.

D1 overhead press is estimated from the existing bench-press reference model using a 0.64 press-to-bench ratio with additional variation. The ratio is anchored to self-selected lifter standards, not paired measurements in D1 athletes. The test pages distinguish these assumptions from source findings.

D1 Pendlay rows use the bench model with assumed row-to-bench ratios of 0.90 for men and 1.00 for women, plus additional variation. D1 weighted pull-ups use rough total-load estimates centered at 122 kg for men and 75 kg for women, with standard deviations of 28 and 17 kg. Those are model parameters, not observed D1 averages. Pull-up studies in football players, mixed-sex swimmers and adjacent trained groups inform the estimates; the women's comparison has less direct support. Pull-ups have no sport-specific D1 comparison.

D1 grip combines published baseball, wrestling, tennis, distance-running and gymnastics results where available, with modeled distributions for unmeasured sports. Hand choice and testing posture differ between studies, so even the sport estimates are approximate. The population grip model remains based on the NHANES survey.

How each group of tests was built

Grip strength — measured

We computed survey-weighted percentiles directly from the public NHANES 2011-2014 data files (about 8,900 US adults aged 18-69), and validated the result against two published analyses of the same survey (every cell within 1 kg). Grip is adjusted for height, not bodyweight: in these data grip scales with height at an exponent of about 1.2-1.4 and with body mass at only about 0.2, which matches the published recommendation to normalise grip by height squared.

Barbell lifts and pull-ups — modeled

No one has tested a one-rep max on a random sample of adults, and no one will. Each lift is modeled as a blend of three groups: untrained adults, recreationally trained lifters, and serious lifters. The untrained distributions come from the baseline measurements of untrained volunteers in training studies (corrected downward, because healthy volunteers out-lift the whole population); the trained distributions come from large logged-lift datasets and tested gym-goers. For ages 18-29 we assume 75% of men and 84% of women are untrained with a barbell,16% / 12% recreationally trained and9% / 4% serious — set from industry participation and time-use surveys rather than the much higher self-reported "I do strength training" rates. Overhead press and barbell row have no direct data at all and are derived from the bench press model using well-supported lift ratios. For pull-ups, the share of adults who can do one strict rep (about 60% of young men and 8% of young women) is an informed assumption and the single most assumption-driven number on the site.

Vertical jump — derived; broad jump, triple hop, 5-10-5 and agility T-test — modeled

Vertical jump uses the Canadian Health Measures Survey, the only nationally representative adult jump dataset we found (n = 5,188), shifted up 2-3 cm because jump-and-reach reads higher than the survey's force plate. Broad jump is scaled from Japan's national fitness survey. We found no representative general-population norms for the triple hop or 5-10-5 shuttle; they are modeled from samples of recreationally active adults, with the general population set below them. Agility times use the protocol stated for each test; timing and course differences are reflected in each model's uncertainty.

The agility T-test uses a study of 304 college-aged adults, reweighted toward lower sport participation for the adult comparison. Its age curve is borrowed from the modeled shuttle curve. D1 basketball, volleyball and soccer cohorts inform the athlete pool, with other sports modeled explicitly. These estimates are least secure for older adults and sports without direct testing.

1.5-mile run and historical mile results — derived and modeled

No representative sample of adults has ever been timed over a mile. We convert laboratory VO₂max percentiles (FRIEND registry) to a sustainable mile pace, calibrated so the result reproduces the measured mile times of US Army recruits at entry (within about 4%), and cross-checked against Finnish conscript data (n = 627,000) and the parkrun 5k median. Untrained people run about 20% slower than standard equations predict from their VO₂max, so uncalibrated conversions — including the widely reprinted Cooper 1.5-mile "norms", which are treadmill estimates rather than run data — are too fast. D1 distance runners come from the complete 2024 D1 1500 m performance lists (about 3,000 athletes per sex); other sports are converted from published VO₂max testing by sport.

Age and body size

Where real age-stratified data exists (grip, vertical jump, VO₂max) we use it directly. Elsewhere an age factor is applied: maximal strength is modeled at 97%, 91%, 82% and 72% of young-adult values for men in their 30s, 40s, 50s and 60s. Norms are interpolated smoothly by age within the supported range. Most population models cover ages 18–69; sprint and 400 m have narrower ranges. Outside a model’s age range, we leave the population comparison empty rather than reuse another age group. D1 comparisons remain unadjusted for age.

For the general-population comparison, lifts are re-expressed at the US median bodyweight for your sex before ranking: strength rises with body mass at roughly the ⅔ power, so a lighter lifter gets credit. Above the median bodyweight the exponent is halved, because in the general population most extra mass is not muscle. Jumping, hopping, sprinting and running are not size-adjusted — the evidence shows jump height is essentially independent of body size.

The overall rank

Each percentile is converted to a z-score. Scores are averaged within each of four domains — strength, power, agility, endurance — and then across domains, so six barbell lifts can't outvote the endurance category. Vertical and broad jump share one slot, so only the selected result counts in the calculator. The main battery requests a 1.5-mile run. A saved mile result keeps its original distance and its own comparison, and fills that same aerobic slot until a 1.5-mile time is entered.

Turning that average into a percentile needs one more ingredient: how strongly the tests correlate. If they were independent, being 85th percentile in everything would be a one-in-thousands profile; in reality fitness traits travel together. We use an average correlation of about 0.55 within a domain and0.3-0.6 between domains for the general population, drawn from published military and athlete test batteries. The 5-10-5 and agility T-test use a higher estimated correlation of 0.85 in both comparisons: they overlap heavily, but their sprint, shuffle and backward-running patterns are distinct enough to provide some additional evidence. Combining both therefore refines the Agility score without giving Agility more weight in the overall rank. Among pooled D1 athletes strength and endurance are negatively correlated (-0.3), because sport selection produces linemen and distance runners, not people who are both.

Two consequences are worth knowing. Breadth is rare: a uniformly good profile ranks higher overall than any single result in it. And the mirror image holds — being below average at everything is also rarer than being below average at one thing, so an overall rank can be lower than every individual score. No dataset exists of people measured on this whole battery, so the overall rank remains an estimate even after minimum coverage is met.

Minimum coverage before a score appears

Individual test percentiles appear immediately. A category percentile requires more than half of its distinct ranked test types with reference data for the selected age, sex and comparison, except Agility, where either supported change-of-direction test is enough. Strength also requires at least one upper-body and one lower-body measurement. Vertical and broad jump share one test type. Historical mile and 1.5-mile results also share one aerobic slot.

Where all age comparisons are available, the general-population minimum is four of seven strength tests, two of three power types (jump, triple hop, sprint), either the 5-10-5 or agility T-test, and both endurance types (1.5-mile and 400 m). Tests without an age-supported estimate are removed from that comparison’s available count. The overall percentile appears only when all four categories meet their minimums.

D1 coverage is calculated separately using the reference data available for the selected sport and sex. Categories without suitable reference coverage stay unavailable. Unranked performance tests and health checks never unlock a percentile.

This is a product rule to avoid presenting a narrow set of measurements as a complete profile, not a statistically validated accuracy threshold. More coverage does not remove uncertainty in the underlying reference models. Compare progress using the same tests.

Floors

The 7 pass/fail standards (Nordic curl, Copenhagen plank, calf raise, suitcase carry, overhead squat, knee-to-wall, sit-to-floor) never affect the rank. They are floors to clear and maintain. The carry load and plank time are scaled to sex, age and bodyweight so the standard is comparably hard for everyone; the others are fixed criteria. The pass rates shown are estimates — real distribution data exists for the calf raise, knee-to-wall and sitting-rising tests, and the rest are modeled.

Known weaknesses

Sources

Each test page lists its own sources. Sources for the floors:

Pooling weights: NCAA Sports Sponsorship and Participation Rates Report, 2024-25. US body size: Fryar CD, et al. Anthropometric reference data for children and adults: United States, 2015-2018. Vital Health Stat 3(46). 2021.