Statistics for clinical research

Statistics you can put your name on.

The most evidence-based profession on earth produces evidence it cannot read.

Average score when 277 internal medicine residents were tested on the biostatistics in the literature they cite.

Windish et al., JAMA 2007

Of the residents most confident in their interpretation of statistical significance, every single one answered incorrectly.

Windish et al., JAMA 2007

Every result comes from a fixed, deterministic statistical engine, never a language model.

That is the gap we built StatsPlease to close. Upload your dataset and get the right test, a finished baseline-characteristics table, and Methods and Results in journal style, every number computed, and yours to recompute.

Free during open access. No credit card.

Results · Vanderbilt diabetes study (n = 390)

A finished result, from a real dataset

Stabilized serum glucose was strongly correlated with glycated hemoglobin (HbA1c), r = 0.75, P < .001: as glucose rose, so did HbA1c.

Baseline characteristics by glycemic status
MeasureHbA1c < 7%HbA1c ≥ 7%P
n33060
Age, y44.7 (16.1)58.4 (13.1)<.001
Glucose, mg/dL91.6 (26.9)194.2 (77.4)<.001

Mean (SD). Mann-Whitney U; groups split at the clinical HbA1c cutoff of 7.0%.

And a null result, reported the same way

HbA1c did not differ between men and women (median 4.90 vs 4.79%; P = .225). Not every comparison is significant, and StatsPlease says so.

A finished result, from real data Tap to read the worked example

See it work

From spreadsheet to result, in three steps.

Upload your data, pick the tests you need, and read a result you can check. Switch the tabs to walk through the flow.

Step 1: Upload your spreadsheet.

Drop in a CSV or Excel file. StatsPlease reads your columns and works out what each one is.

Illustrative preview of an uploaded clinical dataset
IDAgeGlucoseHbA1cGroup
00154965.1Control
002611837.4Diabetes
00347884.9Control
004592018.0Diabetes

Free during open access. No credit card.

The worked example

One dataset. The whole manuscript section.

This is the public Vanderbilt diabetes study, the same starter dataset in every new account. Here is exactly what StatsPlease wrote from it: an Abstract, a Results paragraph, baseline characteristics, and a Methods note. The numbers are computed from the data. The wording is drafted for you to check and edit.

Glycemic control in the Vanderbilt diabetes cohort

Worked example · n = 390 · computed from diabetes.csv

Abstract

We examined glycemic control in 390 adults from a community diabetes screening study. Participants were grouped by glycated hemoglobin (HbA1c) at the clinical cutoff of 7.0%: 60 met the threshold for diabetes and 330 did not. Stabilized serum glucose was strongly correlated with HbA1c (r = 0.75, P < .001). Participants at or above the HbA1c threshold were older (mean 58.4 vs 44.7 years) and had higher stabilized glucose (194.2 vs 91.6 mg/dL) and total cholesterol (228.6 vs 203.4 mg/dL), all P < .001. HbA1c did not differ by sex (P = .225). Glucose and HbA1c track closely in this cohort, while several routine measures separate the two glycemic groups.

Results

Of 390 participants with a recorded HbA1c, 60 (15.4%) were at or above the 7.0% cutoff for diabetes. Stabilized serum glucose was strongly and positively correlated with HbA1c (Pearson r = 0.75, 95% CI 0.70 to 0.79, P < .001).

Compared with participants below the cutoff, those at or above it were older (58.4 [13.1] vs 44.7 [16.1] years), had higher stabilized glucose (194.2 [77.4] vs 91.6 [26.9] mg/dL), higher total cholesterol (228.6 [56.5] vs 203.4 [41.1] mg/dL), and higher systolic blood pressure (147.8 [20.5] vs 135.2 [22.9] mm Hg); each comparison was significant at P < .001 by the Mann-Whitney U test. HDL cholesterol was modestly lower in the higher-HbA1c group (45.3 [16.9] vs 51.2 [17.3] mg/dL, P = .005).

Not every comparison reached significance: HbA1c itself did not differ between men and women (median 4.90 vs 4.79%, P = .225).

Baseline characteristics by glycemic status.
Characteristic HbA1c < 7.0% (n = 330) HbA1c ≥ 7.0% (n = 60) P
Age, years44.7 (16.1)58.4 (13.1)<.001
Stabilized glucose, mg/dL91.6 (26.9)194.2 (77.4)<.001
Total cholesterol, mg/dL203.4 (41.1)228.6 (56.5)<.001
HDL cholesterol, mg/dL51.2 (17.3)45.3 (16.9).005
Systolic BP, mm Hg135.2 (22.9)147.8 (20.5)<.001

Values are mean (SD). P values from the Mann-Whitney U test. Groups defined by the clinical HbA1c cutoff of 7.0%.

Methods

Continuous variables are summarized as mean (standard deviation). Distributions were assessed for normality with the Shapiro-Wilk test and for equal variance with Levene's test. Because the continuous measures departed from normality, two-group comparisons used the Mann-Whitney U test, and the association between stabilized glucose and HbA1c used the Pearson correlation coefficient with a 95% confidence interval. A two-sided P value below .05 was treated as significant.

The test for each comparison was selected automatically from the data through a fixed decision tree, and every value was computed directly from the dataset, not generated by ChatGPT or another large language model. The analysis reproduces exactly on a rerun.

Shown in AMA style — switch to APA in the app.

Computed from the public Vanderbilt diabetes study. Download the data and recompute it in SPSS.

Reproducibility

Don't trust us. Recompute us.

Each result below is a real analysis StatsPlease ran on a published clinical dataset. The test, the numbers, and the data are all here. Download any one, run it in SPSS, and check every value. Some results are strong and some are small, exactly as the data is.

Pearson correlation

Vanderbilt diabetes study · n = 390

Stabilized glucose and HbA1c were strongly correlated; r = 0.75; P < .001.

Both measures continuous and approximately linear, so Pearson's correlation.

Download the data (CSV)

Mann-Whitney U

Heart-failure cohort · n = 299

Serum creatinine was higher in patients who died than in survivors (median 1.30 vs 1.00 mg/dL); U = 14190; P < .001; r = 0.46.

Creatinine was non-normal with unequal variances, so Mann-Whitney U.

Download the data (CSV)

Mann-Whitney U

Melanoma survival study · n = 205

Tumour thickness was greater in ulcerated melanomas (median 3.54 vs 1.29 mm); U = 8520; P < .001; r = 0.65.

Thickness was non-normal in both groups, so Mann-Whitney U.

Download the data (CSV)

How it compares

How StatsPlease compares

The everyday tools researchers reach for, and where each one leaves you.

A comparison of SPSS, R, ChatGPT and other LLMs, and StatsPlease across seven research capabilities.
Capability SPSS R ChatGPT & other LLMs StatsPlease
Numbers you can independently check
Correct test selected automatically
Drafts Methods & Results for you
No coding or stats knowledge needed
Publication-ready output
Reproducible in SPSS
Verified AMA/APA reporting

ChatGPT and other LLMs can produce statistics that read plausibly but cannot be independently checked or reproduced.

Coming from SPSS? See the full side-by-side: StatsPlease as an SPSS alternative.

Free during open access. No credit card.

Scope

What it does not handle

StatsPlease covers the everyday analyses of clinical and epidemiological research. It is deliberately not built for the methods below. For these, a biostatistician remains the right choice.

  • Survival analysis (Kaplan-Meier, Cox regression)
  • Longitudinal and mixed-effects models
  • Complex survey designs with weighting and clustering
  • Bayesian methods

StatsPlease is for the routine analyses researchers currently run in SPSS or send to a statistician for standard processing.

Questions

Common questions

Are the numbers computed or invented?

The numbers are computed. Every test statistic, p-value, effect size, and confidence interval is calculated directly from your data through a fixed, transparent decision tree, the same way an analyst would compute them by hand. They are not generated by ChatGPT or another large language model, so they reproduce exactly and hold up in review. The Methods and Results wording around those numbers is drafted for you in journal style, then yours to check and edit. So the figures are computed and the prose is drafted: nothing about a number is guessed.

What do I actually get?

A complete analysis of your dataset: the statistical tests, publication-quality figures, a baseline-characteristics table, and Methods and Results paragraphs drafted in journal style. The output is built to be pasted into a manuscript, not exported and reworked.

Is it really free?

Yes. StatsPlease is in open access and every feature is currently free, with no credit card required.

How do I know the statistics are correct?

Core statistics are checked against reference values computed independently of StatsPlease's own code: R outputs, textbook closed forms, and NIST-certified reference datasets. We also publish worked examples on real, widely cited clinical datasets, each with the chart, the result, and the data, so you can download any of them and reproduce every number yourself. Run your own analysis twice and you will get identical numbers each time. Read the validation summary.

Which analyses does StatsPlease run?

The everyday tests of clinical research: two-group and multi-group comparisons, paired analyses, correlations, regression, and categorical comparisons, in parametric and nonparametric forms. StatsPlease checks the assumptions and selects an appropriate test for your data; see our guide on choosing between Mann-Whitney U and the t-test.

What does StatsPlease not handle?

StatsPlease handles the everyday analyses of clinical and epidemiological research: group comparisons, correlations, regression, and categorical tests, in parametric and nonparametric forms. It is not designed for survival analysis, longitudinal mixed models, complex survey designs, or Bayesian methods. For those, a biostatistician remains the right choice. StatsPlease is built for the routine analyses researchers currently run in SPSS or send to a statistician for standard processing.

Do I need to install anything?

No. StatsPlease runs in your browser. Create a free account, upload a dataset, and your first written results arrive in minutes.

What does the AI actually do, and can I turn it off?

After our deterministic engine computes your results, an "AI Polish" toggle in the Analyze flow sends the draft Results wording (not your dataset) to our AI sub-processor, Anthropic's Claude API, to smooth the prose. The AI never computes a number: every statistic, p-value, effect size, and confidence interval comes from our own engine, and the polished text is checked against those original values before you see it. The toggle is on by default, and one click turns it off entirely: same computed numbers, fully deterministic AMA wording, zero AI involved for that analysis.

Free during open access. No credit card.

Why we built this

Statistics is the last manual process left.

Medicine built a machine to read an ECG. We built StatsPlease for the same reason: the number in your Results section will be cited by other researchers, used to support clinical decisions, and presented to ethics committees. It should be correct, and you should understand why.

Read the manifesto →

Blog

Guides for clinical researchers

Plain-language guides on choosing the right test and reporting results for publication.

Browse all guides on the blog →

Statistics you can put your name on.

Start free