Pearson correlation
Vanderbilt diabetes study · n = 390
Stabilized glucose and HbA1c were strongly correlated; r = 0.75; P < .001.
Both measures continuous and approximately linear, so Pearson's correlation.
Download the data (CSV)Statistics for clinical research
The most evidence-based profession on earth produces evidence it cannot read.
Average score when 277 internal medicine residents were tested on the biostatistics in the literature they cite.
Of the residents most confident in their interpretation of statistical significance, every single one answered incorrectly.
Every result comes from a fixed, deterministic statistical engine, never a language model.
That is the gap we built StatsPlease to close. Upload your dataset and get the right test, a finished baseline-characteristics table, and Methods and Results in journal style, every number computed, and yours to recompute.
Free during open access. No credit card.
Results · Vanderbilt diabetes study (n = 390)
Stabilized serum glucose was strongly correlated with glycated hemoglobin (HbA1c), r = 0.75, P < .001: as glucose rose, so did HbA1c.
| Measure | HbA1c < 7% | HbA1c ≥ 7% | P |
|---|---|---|---|
| n | 330 | 60 | |
| Age, y | 44.7 (16.1) | 58.4 (13.1) | <.001 |
| Glucose, mg/dL | 91.6 (26.9) | 194.2 (77.4) | <.001 |
Mean (SD). Mann-Whitney U; groups split at the clinical HbA1c cutoff of 7.0%.
And a null result, reported the same way
HbA1c did not differ between men and women (median 4.90 vs 4.79%; P = .225). Not every comparison is significant, and StatsPlease says so.
See it work
Upload your data, pick the tests you need, and read a result you can check. Switch the tabs to walk through the flow.
Drop in a CSV or Excel file. StatsPlease reads your columns and works out what each one is.
| ID | Age | Glucose | HbA1c | Group |
|---|---|---|---|---|
| 001 | 54 | 96 | 5.1 | Control |
| 002 | 61 | 183 | 7.4 | Diabetes |
| 003 | 47 | 88 | 4.9 | Control |
| 004 | 59 | 201 | 8.0 | Diabetes |
Click the tests you need, or let our full-analysis algorithm generate every relevant result.
StatsPlease checks the assumptions and picks the right form of each test for your data.
Here is a correlation between stabilized glucose and HbA1c, reported in journal style.
| Comparison | n | r | 95% CI | P |
|---|---|---|---|---|
| Glucose vs HbA1c | 390 | 0.75 | 0.70 to 0.79 | < .001 |
Representative output. The numbers shown here are illustrative.
Free during open access. No credit card.
The worked example
This is the public Vanderbilt diabetes study, the same starter dataset in every new account. Here is exactly what StatsPlease wrote from it: an Abstract, a Results paragraph, baseline characteristics, and a Methods note. The numbers are computed from the data. The wording is drafted for you to check and edit.
We examined glycemic control in 390 adults from a community diabetes screening study. Participants were grouped by glycated hemoglobin (HbA1c) at the clinical cutoff of 7.0%: 60 met the threshold for diabetes and 330 did not. Stabilized serum glucose was strongly correlated with HbA1c (r = 0.75, P < .001). Participants at or above the HbA1c threshold were older (mean 58.4 vs 44.7 years) and had higher stabilized glucose (194.2 vs 91.6 mg/dL) and total cholesterol (228.6 vs 203.4 mg/dL), all P < .001. HbA1c did not differ by sex (P = .225). Glucose and HbA1c track closely in this cohort, while several routine measures separate the two glycemic groups.
Of 390 participants with a recorded HbA1c, 60 (15.4%) were at or above the 7.0% cutoff for diabetes. Stabilized serum glucose was strongly and positively correlated with HbA1c (Pearson r = 0.75, 95% CI 0.70 to 0.79, P < .001).
Compared with participants below the cutoff, those at or above it were older (58.4 [13.1] vs 44.7 [16.1] years), had higher stabilized glucose (194.2 [77.4] vs 91.6 [26.9] mg/dL), higher total cholesterol (228.6 [56.5] vs 203.4 [41.1] mg/dL), and higher systolic blood pressure (147.8 [20.5] vs 135.2 [22.9] mm Hg); each comparison was significant at P < .001 by the Mann-Whitney U test. HDL cholesterol was modestly lower in the higher-HbA1c group (45.3 [16.9] vs 51.2 [17.3] mg/dL, P = .005).
Not every comparison reached significance: HbA1c itself did not differ between men and women (median 4.90 vs 4.79%, P = .225).
| Characteristic | HbA1c < 7.0% (n = 330) | HbA1c ≥ 7.0% (n = 60) | P |
|---|---|---|---|
| Age, years | 44.7 (16.1) | 58.4 (13.1) | <.001 |
| Stabilized glucose, mg/dL | 91.6 (26.9) | 194.2 (77.4) | <.001 |
| Total cholesterol, mg/dL | 203.4 (41.1) | 228.6 (56.5) | <.001 |
| HDL cholesterol, mg/dL | 51.2 (17.3) | 45.3 (16.9) | .005 |
| Systolic BP, mm Hg | 135.2 (22.9) | 147.8 (20.5) | <.001 |
Values are mean (SD). P values from the Mann-Whitney U test. Groups defined by the clinical HbA1c cutoff of 7.0%.
Continuous variables are summarized as mean (standard deviation). Distributions were assessed for normality with the Shapiro-Wilk test and for equal variance with Levene's test. Because the continuous measures departed from normality, two-group comparisons used the Mann-Whitney U test, and the association between stabilized glucose and HbA1c used the Pearson correlation coefficient with a 95% confidence interval. A two-sided P value below .05 was treated as significant.
The test for each comparison was selected automatically from the data through a fixed decision tree, and every value was computed directly from the dataset, not generated by ChatGPT or another large language model. The analysis reproduces exactly on a rerun.
Shown in AMA style — switch to APA in the app.
Computed from the public Vanderbilt diabetes study. Download the data and recompute it in SPSS.
Reproducibility
Each result below is a real analysis StatsPlease ran on a published clinical dataset. The test, the numbers, and the data are all here. Download any one, run it in SPSS, and check every value. Some results are strong and some are small, exactly as the data is.
Pearson correlation
Vanderbilt diabetes study · n = 390
Stabilized glucose and HbA1c were strongly correlated; r = 0.75; P < .001.
Both measures continuous and approximately linear, so Pearson's correlation.
Download the data (CSV)Mann-Whitney U
Heart-failure cohort · n = 299
Serum creatinine was higher in patients who died than in survivors (median 1.30 vs 1.00 mg/dL); U = 14190; P < .001; r = 0.46.
Creatinine was non-normal with unequal variances, so Mann-Whitney U.
Download the data (CSV)Mann-Whitney U
Melanoma survival study · n = 205
Tumour thickness was greater in ulcerated melanomas (median 3.54 vs 1.29 mm); U = 8520; P < .001; r = 0.65.
Thickness was non-normal in both groups, so Mann-Whitney U.
Download the data (CSV)How it compares
The everyday tools researchers reach for, and where each one leaves you.
| Capability | SPSS | R | ChatGPT & other LLMs | StatsPlease |
|---|---|---|---|---|
| Numbers you can independently check | ✓ | ✓ | ✗ | ✓ |
| Correct test selected automatically | ✗ | ✗ | ⚠ | ✓ |
| Drafts Methods & Results for you | ✗ | ✗ | ✓ | ✓ |
| No coding or stats knowledge needed | ✗ | ✗ | ✓ | ✓ |
| Publication-ready output | ✗ | ✗ | ⚠ | ✓ |
| Reproducible in SPSS | ✓ | ✓ | ✗ | ✓ |
| Verified AMA/APA reporting | ✗ | ✗ | ✗ | ✓ |
ChatGPT and other LLMs can produce statistics that read plausibly but cannot be independently checked or reproduced.
Coming from SPSS? See the full side-by-side: StatsPlease as an SPSS alternative.
Free during open access. No credit card.
Scope
StatsPlease covers the everyday analyses of clinical and epidemiological research. It is deliberately not built for the methods below. For these, a biostatistician remains the right choice.
StatsPlease is for the routine analyses researchers currently run in SPSS or send to a statistician for standard processing.
Questions
The numbers are computed. Every test statistic, p-value, effect size, and confidence interval is calculated directly from your data through a fixed, transparent decision tree, the same way an analyst would compute them by hand. They are not generated by ChatGPT or another large language model, so they reproduce exactly and hold up in review. The Methods and Results wording around those numbers is drafted for you in journal style, then yours to check and edit. So the figures are computed and the prose is drafted: nothing about a number is guessed.
A complete analysis of your dataset: the statistical tests, publication-quality figures, a baseline-characteristics table, and Methods and Results paragraphs drafted in journal style. The output is built to be pasted into a manuscript, not exported and reworked.
Yes. StatsPlease is in open access and every feature is currently free, with no credit card required.
Core statistics are checked against reference values computed independently of StatsPlease's own code: R outputs, textbook closed forms, and NIST-certified reference datasets. We also publish worked examples on real, widely cited clinical datasets, each with the chart, the result, and the data, so you can download any of them and reproduce every number yourself. Run your own analysis twice and you will get identical numbers each time. Read the validation summary.
The everyday tests of clinical research: two-group and multi-group comparisons, paired analyses, correlations, regression, and categorical comparisons, in parametric and nonparametric forms. StatsPlease checks the assumptions and selects an appropriate test for your data; see our guide on choosing between Mann-Whitney U and the t-test.
StatsPlease handles the everyday analyses of clinical and epidemiological research: group comparisons, correlations, regression, and categorical tests, in parametric and nonparametric forms. It is not designed for survival analysis, longitudinal mixed models, complex survey designs, or Bayesian methods. For those, a biostatistician remains the right choice. StatsPlease is built for the routine analyses researchers currently run in SPSS or send to a statistician for standard processing.
No. StatsPlease runs in your browser. Create a free account, upload a dataset, and your first written results arrive in minutes.
After our deterministic engine computes your results, an "AI Polish" toggle in the Analyze flow sends the draft Results wording (not your dataset) to our AI sub-processor, Anthropic's Claude API, to smooth the prose. The AI never computes a number: every statistic, p-value, effect size, and confidence interval comes from our own engine, and the polished text is checked against those original values before you see it. The toggle is on by default, and one click turns it off entirely: same computed numbers, fully deterministic AMA wording, zero AI involved for that analysis.
Free during open access. No credit card.
Blog
Plain-language guides on choosing the right test and reporting results for publication.
Test selection
Mann-Whitney U vs t-test: which to use
How to choose between the two most common comparison tests, with normality testing and reporting guidance.
Read guide →Start here
The profession that can’t read its own evidence
EBM taught doctors to cite statistics, not read them. Five numbers that explain the problem, and why we built StatsPlease.
Read the manifesto →Interpretation
The one sentence about statistical significance most papers get wrong, and what a P-value can and cannot tell you.
Read guide →Statistics you can put your name on.
Start free