How this instrument works
The San Antonio Diabetes Prediction Model is a logistic-regression risk score built by Stern, Williams, and Haffner from the San Antonio Heart Study and published in Annals of Internal Medicine in 2002. It combines age, sex, Mexican American ethnicity, fasting plasma glucose, systolic blood pressure, HDL cholesterol, body mass index, and family history of diabetes into a single weighted sum (the logit), which is then converted into a predicted probability of developing type 2 diabetes over the following 7.5 years. It is an early example of a diabetes risk equation that folds a fasting glucose measurement directly into the regression rather than relying only on anthropometric and demographic factors.
On sourcing: the 2002 primary paper sits behind a journal paywall and could not be read directly while building this page. Instead, the intercept (−13.415) and all eight regression coefficients were independently cross-checked against two separate peer-reviewed papers that each cite Stern 2002 directly and reproduce the identical intercept and coefficients to three decimal places — Mann et al., Multi-Ethnic Study of Atherosclerosis (MESA), American Journal of Epidemiology, 2010; and Lacy et al., CARDIA Study, Diabetes Care, 2016. Two independent teams reproducing the same numbers from the same source gives good confidence the formula below is accurate, even without direct access to the original.
A scope limitation deserves top billing, not a footnote: the model was derived exclusively from Mexican American and non-Hispanic white participants of the San Antonio Heart Study — the derivation cohort included zero Black, zero Asian, and zero other-Hispanic-subgroup participants. Later validation work has found real miscalibration when the equation is applied outside that population. A Tehran cohort study found the model overestimated diabetes risk by roughly 111% before recalibration. A Japanese American cohort study found inconsistent performance across age groups. A 2016 secondary analysis of the decades-old CARDIA cohort (Lacy et al.) tested this model — alongside two other diabetes risk equations — in Black and white participants, precisely because none were represented when the San Antonio model was built, and found the same white-versus-Black discrimination gap in this model as in the other two. Even in a population that overlaps substantially with the original derivation cohort, the model's own MESA validation study found it overestimated absolute risk across every predicted-risk level before recalibration — treat the percentage this calculator returns as a relative indicator of risk rather than a precise probability.
This is a research-grade risk-prediction tool intended for clinicians and researchers estimating population-level diabetes risk, not a diagnostic instrument and not a substitute for an actual glucose test. It requires a real fasting plasma glucose value from a blood draw as one of its eight inputs — it is not a screening tool that skips laboratory work, and a result here should prompt a conversation with a clinician about testing and prevention, not stand in for one.
- Enter age in years (age) — the underlying cohort was a mid-life adult population, not children or the very elderly.
- Select sex (female) — Yes adds a fixed 0.661 to the logit regardless of any other input.
- Select Mexican American ethnicity (mexAm) — Yes adds 0.412; read the scope limitations above before applying this field to other ethnic groups.
- Enter fasting plasma glucose in mg/dL (glucose) from an actual lab blood draw, not an estimate or a random (non-fasting) reading.
- Enter systolic blood pressure (sbp), HDL cholesterol (hdl), height and weight (height, weight — used to compute bmi), and family history of diabetes in a parent or sibling (familyHistory).
- Read the computed body mass index (bmi), the logistic linear predictor (logit), and the final predicted 7.5-year diabetes risk (riskPercent).
Worked examples across the risk range
No worked numerical example for the San Antonio Diabetes Prediction Model appears in Stern's 2002 paper or in any secondary source found during research for this page, so every case below is self-built and checked only for the correct direction of each risk factor's effect, not against a published figure. A 30-year-old non-Hispanic-white man (female = No, mexAm = No) with fasting glucose 85 mg/dL, systolic blood pressure 110 mmHg, HDL 55 mg/dL, height 1.75 m, weight 70 kg (BMI 22.9), and no family history of diabetes produces a logit of −4.4250, which works out to a predicted 7.5-year diabetes risk of about 1.2%.
A 50-year-old Mexican American woman (female = Yes, mexAm = Yes) with borderline-elevated glucose (105 mg/dL), mild hypertension (SBP 130 mmHg), HDL 45 mg/dL, height 1.60 m, weight 75 kg (BMI 29.3), and a family history of diabetes lands in the middle of the range: a logit of 0.4698 and a predicted risk of about 61.5%, reflecting how sex, ethnicity, and family history each add fixed points to the weighted sum.
At the higher end, a 65-year-old Mexican American man (female = No, mexAm = Yes) with impaired fasting glucose (125 mg/dL), hypertension (SBP 150 mmHg), low HDL (35 mg/dL), height 1.70 m, weight 95 kg (BMI 32.9), and a family history of diabetes produces a logit of 2.8090, which works out to a predicted risk of about 94.3% — showing how strongly these eight factors compound once several of them point the same direction at once.
Questions
What does the predicted risk percentage actually mean?
It is the model's estimated probability that a person with the entered characteristics develops type 2 diabetes within 7.5 years, based on the rate observed among similar participants in the San Antonio Heart Study derivation cohort. It is a population-based probability, not a certainty and not a diagnosis — a low percentage does not rule out diabetes, and a high one does not confirm it. It is meant to inform a conversation with a clinician about testing and prevention, not to replace one.
Who was this model actually built and tested on?
The San Antonio Diabetes Prediction Model was derived exclusively from Mexican American and non-Hispanic white participants of the San Antonio Heart Study. No Black, Asian, or other Hispanic-subgroup participants were included in the derivation cohort. That matters directly for how far the coefficients — including the fixed +0.412 added for Mexican American ethnicity — can reasonably be expected to generalize to people outside those two groups.
Does this model work well outside its original San Antonio cohort?
Inconsistently, even within a broadly similar population. A Tehran validation study found the model overestimated diabetes risk by roughly 111% before recalibration. A Japanese American study found performance varied inconsistently by age. A 2016 secondary analysis of the CARDIA cohort (Lacy et al.) tested this model, alongside two other risk equations, in Black Americans, a group with zero representation in the derivation cohort, and found meaningful racial differences in accuracy for all three. Even the MESA study, in a population overlapping the original cohort, found this model overestimated absolute risk before recalibration. Treat the output as relative, not precise.
Why does the calculator need a fasting glucose blood draw?
Fasting plasma glucose is one of the eight regression inputs in the published formula itself, carrying a coefficient of 0.079 per mg/dL, so the model cannot run without an actual measured value. This is a research-grade risk-prediction tool that expects a real laboratory result, not a no-blood-test screening shortcut — a random or estimated glucose value will not produce a reliable output.
How were the coefficients verified if the original 2002 paper is paywalled?
The intercept and all eight coefficients were cross-checked against two independent peer-reviewed papers that each cite Stern 2002 directly and reproduce the same intercept (−13.415) and coefficients to three decimal places: Mann et al. in the Multi-Ethnic Study of Atherosclerosis (MESA), American Journal of Epidemiology, 2010, and Lacy et al. in the CARDIA Study, Diabetes Care, 2016. Two independent research groups reproducing identical figures from the same source is strong indirect confirmation, even without direct access to the original text.
Can this calculator diagnose diabetes or prediabetes?
No. It performs the published logistic-regression arithmetic on the values entered and returns a probability estimate; it does not diagnose anything. Diagnosing diabetes or prediabetes requires actual clinical testing — fasting plasma glucose, an oral glucose tolerance test, or hemoglobin A1c — interpreted by a clinician against recognized diagnostic thresholds, not a risk score.
Why do sex and ethnicity change the result by a fixed amount?
In this model, female sex adds a fixed 0.661 to the logit and Mexican American ethnicity adds a fixed 0.412, regardless of any other input — unlike some other risk models, there is no interaction term that changes those effects based on age or other factors. Both coefficients were estimated from the San Antonio Heart Study cohort specifically, which is also why applying the ethnicity field to groups outside that derivation population carries the scope caution described above.
What counts as a qualifying family history of diabetes?
The model defines family history as diabetes in a parent or a sibling — that is the definition used in the original derivation and the one this calculator's familyHistory field reflects. Set it to Yes only if a parent or sibling has been diagnosed with diabetes; more distant relatives (grandparents, aunts, uncles, cousins) were not part of the original definition and are not what this field is asking about.
References
- Stern MP, Williams K, Haffner SM — Ann Intern Med. 2002;136(8):575-581 (PMID 11955025)
- Mann DM, Bertoni AG, Shimbo D, et al. — MESA cohort, Am J Epidemiol. 2010;171(9):980-988 (PMC2877477)
- Lacy ME, Wellenius GA, Carnethon MR, et al. — CARDIA Study, Diabetes Care. 2016;39(2):285-291 (PMC4722943)
- Bozorgmanesh M, Hadaegh F, Zabetian A, Azizi F — Tehran validation study (PMID 20217177)
- McNeely MJ, Boyko EJ, et al. — Japanese Americans, Diabetes Care. 2003;26(3):758-763 (PMID 12610034)
Read this first: This instrument computes a screening figure from population formulas — it is not a diagnosis, and it cannot see the whole picture a clinician can. Use it to inform a conversation, not to replace one.