Cochrane Database of Systematic Reviews (CDSR)
To download only this data file: Cochrane.rds (42 MB)
To download all BEAR datasets, click here.
Description
Trials with results available in CDSR, https://www.cochranelibrary.com/ (Cochrane Collaboration 2025).
Snapshot date: 20 November 2025
Data collection: Using the R package cochrane (Schwab 2024), we combined data from 8,726 CDSR review records into a single file, cdsr_interventions_19nov2025.csv (63 MB).
Data processing: We applied several processing steps and extensive filtering.
During processing, we cleaned up study years and categorised measures (risk ratio, odds ratio, mean difference, etc.) and experimental designs. In particular, we created an “RCT” flag by scanning each review abstract for its inclusion criteria. Most studies in CDSR are RCTs, but some reviews also allowed quasi-experimental studies, so we could not classify many studies. We also classify outcome rows into efficacy, safety, dropouts, and bias based on the comparison, outcome, and subgroup labels.
To calculate effect sizes and z-values for BEAR, we used metafor to calculate standardised mean differences for continuous outcomes and probit-transformed differences for dichotomous outcomes. This puts continuous and binary outcomes on a comparable scale.
Unlike in most of the other datasets in BEAR, we filter CDSR data heavily. There are six important decisions:
- Use only the outcome and comparison from each review coded as “1”, as these are most likely to be the primary outcome and most relevant comparison.
- Keep only rows classified as
efficacy, excluding likely safety, dropout, and bias-related outcomes based on comparison, outcome, and subgroup labels. - Use only studies with continuous and dichotomous outcomes, removing rows where effect estimates are based on instrumental variables or individual patient data.
- Remove rows where the effect measure is unknown (keeping OR, RR, Peto OR, mean difference, standardised mean difference, and risk difference) and rows with no participants.
- Remove count data with no events or a 100% event rate in both arms (see the discussion of zero-event studies).
- Exclude reviews marked as withdrawn in the source RM5 file.
These processing steps are applied in the main dataset, BEAR.rds. Additional rows are still available to researchers in the individual dataset, Cochrane.rds. That file retains withdrawn reviews with a withdrawn indicator (1 for withdrawn, 0 otherwise). Its single doi field identifies the source edition where this could be established from the saved file or an edition audit.
As a result, the distributed BEAR.rds file includes a smaller analysis subset with about 33,500 rows: the first comparison and first outcome from each review, restricted to efficacy outcomes. The processed data/Cochrane.rds file is larger and is available for users who want to make different selection decisions. It contains 760,486 result rows, representing 90,042 review-study pairs across 6,619 Cochrane reviews.
A large proportion of CDSR rows have binary outcomes with no events in either arm or events in all participants in both arms. In the BEAR analysis subset before this exclusion, 2,327 of 24,959 binary rows (9.3%) met one of these conditions: 2,208 rows (8.8%) had no events in either arm and 119 rows (0.5%) had all participants experiencing the event in both arms.
Zero-event studies may be important for some metascientific analyses, but for modelling the distribution of z-values they are unhelpful. With these rows included, the empirical distribution has extra observations at or near zero, and the normal approximation is poor. Because of this peak at zero, including these non-significant studies has the counter-intuitive effect of increasing the estimated degree of selection: the fitted distribution has a steeper drop-off between zero and the conventional significance threshold at 1.96.
Study characteristics:
Sample size and year:
ssis the sum of the two arm totals;yearisstudy.year.Study ID: we use
study.name.Meta-analysis ID: we use the Cochrane review identifier (
id).Topic: we use the Cochrane specialty.
subsetrecords study-data provenance: published, unpublished, sought or mixed.method: We review meta-analysis eligibility criteria to determine which meta-analyses included only randomised trials. This classifies 67% of estimates as RCTs; the remaining 33% could not be classified reliably, although many are likely to be RCTs because Cochrane meta-analyses typically include RCT or quasi-randomised evidence.measure: About 32% of estimates are standardised mean differences and 68% are differences between probit-transformed event proportions.topicis Cochrane specialty andsubsetis source data type (published, unpublished, sought, or mixed);outcome_groupis also available in the processed data.
Model of z-values
| Characteristic | Estimate |
|---|---|
| Probability of significance | 28% |
| Relative probability of publication for |z| < 1.96 | 0.82 |
| Successful replication for |z| > 1.96 | 60% |
| Correct sign for |z| > 1.96 | 97% |
What do these terms mean?
- Probability of significance
- The reported value is the assurance: the proportion of significant results adjusted for publication bias.
- Relative probability of publication
- The relative probability of observing a result below the |z| = 1.96 threshold rather than above it. Values below one indicate lower observation probability below the conventional two-sided significance threshold.
- Successful replication
- The probability that an exact replication has the same sign and a |z| greater than 1.96, conditional on the original result having |z| greater than 1.96.
- Correct sign
- The probability that the observed effect has the same direction as the true effect, conditional on an original result with |z| greater than 1.96.

Data dictionary
One row is one study result within a Cochrane review analysis. Summaries count rows, not distinct reviews.
Review identifiers
| Variable | Definition | Summary |
|---|---|---|
cochrane_id |
Stable CD review identifier from the RM5 filename. | |
doi |
DOI of the RM5 source edition, where established; a checkpoint label is used only when no contrary evidence exists. | |
withdrawn |
Whether the source RM5 review has withdrawal status W; repeated on every study-result row. | 0 733,563; 1 26,923 |
Review characteristics
| Variable | Definition | Summary |
|---|---|---|
specialty |
Cochrane Review Group code for the source edition. | missing 36,034 |
rct |
Whether the matching edition abstract’s eligibility section supports RCT-only classification. | TRUE 505,795; FALSE 187,878; NA 66,813 |
id |
Review identifier supplied by the RM5 parser. |
Analysis identifiers
| Variable | Definition | Summary |
|---|---|---|
comparison.nr |
Comparison number within the review. | |
comparison.name |
Comparison label. | |
comparison.id |
Internal comparison identifier. | |
outcome.nr |
Outcome number within the comparison. | |
outcome.name |
Outcome label. | |
outcome.measure |
Effect measure named for the source outcome. | |
outcome.id |
Internal outcome identifier. | |
outcome.flag |
Source result family. | |
subgroup.nr |
Subgroup number within the outcome. | |
subgroup.name |
Subgroup label. | |
subgroup.id |
Internal subgroup identifier. | missing 248,547 |
Study characteristics
| Variable | Definition | Summary |
|---|---|---|
study.id |
Internal source study identifier. | |
study.name |
Study label used as studyid in BEAR. | |
study.year |
Year parsed from the source study label or year field; implausible years are missing. | missing 27,883 |
study.data_source |
Source status of the study data. | PUB 630,749; MIX 101,642; SOUGHT 17,328; UNPUB 10,767 |
Reported effects
| Variable | Definition | Summary |
|---|---|---|
effect.size |
Effect estimate reported in the RM5 study table. | |
se |
Reported standard error for effect.size. | |
ci.lower |
Reported confidence interval lower limit. | |
ci.upper |
Reported confidence interval upper limit. | |
weight |
Analysis weight reported in the RM5 study table. | |
order |
Source ordering of the study result within the analysis. |
Arm inputs
| Variable | Definition | Summary |
|---|---|---|
events1 |
Event count in arm 1 for dichotomous outcomes. | missing 249,625 |
total1 |
Participant count in arm 1. | |
mean1 |
Mean in arm 1 for continuous outcomes. | missing 510,861 |
sd1 |
Standard deviation in arm 1 for continuous outcomes. | missing 510,861 |
events2 |
Event count in arm 2 for dichotomous outcomes. | missing 249,625 |
total2 |
Participant count in arm 2. | |
mean2 |
Mean in arm 2 for continuous outcomes. | missing 510,861 |
sd2 |
Standard deviation in arm 2 for continuous outcomes. | missing 510,861 |
BEAR calculations
| Variable | Definition | Summary |
|---|---|---|
measure_group |
Broad effect measure category from outcome.measure. | |
outcome_group |
Heuristic outcome category from comparison, outcome and subgroup labels. | efficacy 595,500; safety 100,241; bias 46,480; dropouts 18,265 |
phase |
Unused placeholder for trial phase. | missing 760,486 |
yi |
Recalculated Hedges g or probit difference from arm inputs. | missing 2,611 |
vi |
Sampling variance of yi. | missing 2,611 |
measure |
Broad label for the recalculated effect. | probit 510,861; SMD 249,625 |
measure_detailed |
Detailed label for the recalculated effect. | |
z |
yi divided by the square root of vi. | missing 2,611 |