SCORE
To download only this data file: SCORE_all_claims.rds (232 KB); SCORE_replications.rds (47 KB)
To download all BEAR datasets, click here.
Description
Reference: Tyner et al. (2026).
Research question: investigating the replicability of published claims in the social and behavioural sciences.
Data availability: The SCORE replication package is available at https://osf.io/g5sny/. The package includes replication data as well as reviews of claims made in the included papers.
Data description and source: BEAR includes two SCORE outputs. The source dataset of replications has 548 rows: 274 original claims and 274 matched replications. After BEAR processing and filtering, 267 matched replication rows are included. The source dataset for “all claims” has 3,066 rows for claims from the set of replicated papers (but no replications of these claims), where we select one statistic per claim; after BEAR processing and filtering, 1,942 claim rows are included.
Data processing: For the matched replication output, we used the package’s converted correlation scale where available for original and replication statistics. For the set of all claims, effects are heterogeneous across the source papers.
When several inputs are available for constructing a z-value in the set of all claims, we prefer reported z, then reported t, then coefficient divided by standard error, then signed square-root F for numerator df 1, then a 95% confidence interval with a point estimate, then a two-sided p-value conversion. When the selected statistic only supplies a p-value and no sign can be inferred, the z-value is unsigned. In the text for all claims, it is typical to have multiple statistics backing up a single claim, e.g. “F(1,38) = 3.73, p = .033, partial eta-squared = .16; F(1,39) = 7.28, p = .010, partial eta-squared = .16; F<1”. We created a rule to pick one statistic per claim using agreement with the SCORE classification as significant or non-significant, the statistic’s provenance, exactness, the available effect and standard error, and text order.
Study characteristics:
- Study ID: we use the DOI/source-paper identifier in both outputs.
- Topic: field of study (e.g. “political science”) as classified by the authors.
- The underlying study design is not systematically described.
- Measures: SCORE “claims” dataset measures include regression coefficients, standardised mean differences, correlations, and variance-explained statistics; but the effect measure is unclassified for over 80% of data. SCORE replications use correlations where available.
Model of z-values
This documentation page covers more than one fitted dataset, so the fitted models are shown separately.
| Characteristic | Estimate |
|---|---|
| Probability of significance | 47% |
| Relative probability of publication for |z| < 1.96 | 0.12 |
| Successful replication for |z| > 1.96 | 73% |
| Correct sign for |z| > 1.96 | 99% |
| Characteristic | Estimate |
|---|---|
| Probability of significance | 51% |
| Relative probability of publication for |z| < 1.96 | 0.70 |
| Successful replication for |z| > 1.96 | 78% |
| Correct sign for |z| > 1.96 | 99% |
What do these terms mean?
- Probability of significance
- The reported value is the assurance: the proportion of significant results adjusted for publication bias.
- Relative probability of publication
- The relative probability of observing a result below the |z| = 1.96 threshold rather than above it. Values below one indicate lower observation probability below the conventional two-sided significance threshold.
- Successful replication
- The probability that an exact replication has the same sign and a |z| greater than 1.96, conditional on the original result having |z| greater than 1.96.
- Correct sign
- The probability that the observed effect has the same direction as the true effect, conditional on an original result with |z| greater than 1.96.

