Cochrane Database of Systematic Reviews (CDSR)

Tags:meta analysisdatabaseprimary outcomeRCTHover over the tag text for details.
Domainmedicine & health
Data90k studies in 6.6k Cochrane reviews
Note

To download only this data file: Cochrane.rds (42 MB)

To download all BEAR datasets, click here.

Description

Trials with results available in CDSR, https://www.cochranelibrary.com/ (Cochrane Collaboration 2025).

Snapshot date: 20 November 2025

Data collection: Using the R package cochrane (Schwab 2024), we combined data from 8,726 CDSR review records into a single file, cdsr_interventions_19nov2025.csv (63 MB).

Data processing: We applied several processing steps and extensive filtering.

During processing, we cleaned up study years and categorised measures (risk ratio, odds ratio, mean difference, etc.) and experimental designs. In particular, we created an “RCT” flag by scanning each review abstract for its inclusion criteria. Most studies in CDSR are RCTs, but some reviews also allowed quasi-experimental studies, so we could not classify many studies. We also classify outcome rows into efficacy, safety, dropouts, and bias based on the comparison, outcome, and subgroup labels.

To calculate effect sizes and z-values for BEAR, we used metafor to calculate standardised mean differences for continuous outcomes and probit-transformed differences for dichotomous outcomes. This puts continuous and binary outcomes on a comparable scale.

Unlike in most of the other datasets in BEAR, we filter CDSR data heavily. There are six important decisions:

  1. Use only the outcome and comparison from each review coded as “1”, as these are most likely to be the primary outcome and most relevant comparison.
  2. Keep only rows classified as efficacy, excluding likely safety, dropout, and bias-related outcomes based on comparison, outcome, and subgroup labels.
  3. Use only studies with continuous and dichotomous outcomes, removing rows where effect estimates are based on instrumental variables or individual patient data.
  4. Remove rows where the effect measure is unknown (keeping OR, RR, Peto OR, mean difference, standardised mean difference, and risk difference) and rows with no participants.
  5. Remove count data with no events or a 100% event rate in both arms (see the discussion of zero-event studies).
  6. Exclude reviews marked as withdrawn in the source RM5 file.

These processing steps are applied in the main dataset, BEAR.rds. Additional rows are still available to researchers in the individual dataset, Cochrane.rds. That file retains withdrawn reviews with a withdrawn indicator (1 for withdrawn, 0 otherwise). Its single doi field identifies the source edition where this could be established from the saved file or an edition audit.

As a result, the distributed BEAR.rds file includes a smaller analysis subset with about 33,500 rows: the first comparison and first outcome from each review, restricted to efficacy outcomes. The processed data/Cochrane.rds file is larger and is available for users who want to make different selection decisions. It contains 760,486 result rows, representing 90,042 review-study pairs across 6,619 Cochrane reviews.

A large proportion of CDSR rows have binary outcomes with no events in either arm or events in all participants in both arms. In the BEAR analysis subset before this exclusion, 2,327 of 24,959 binary rows (9.3%) met one of these conditions: 2,208 rows (8.8%) had no events in either arm and 119 rows (0.5%) had all participants experiencing the event in both arms.

Zero-event studies may be important for some metascientific analyses, but for modelling the distribution of z-values they are unhelpful. With these rows included, the empirical distribution has extra observations at or near zero, and the normal approximation is poor. Because of this peak at zero, including these non-significant studies has the counter-intuitive effect of increasing the estimated degree of selection: the fitted distribution has a steeper drop-off between zero and the conventional significance threshold at 1.96.

Study characteristics:

  • Sample size and year: ss is the sum of the two arm totals; year is study.year.

  • Study ID: we use study.name.

  • Meta-analysis ID: we use the Cochrane review identifier (id).

  • Topic: we use the Cochrane specialty.

  • subset records study-data provenance: published, unpublished, sought or mixed.

  • method: We review meta-analysis eligibility criteria to determine which meta-analyses included only randomised trials. This classifies 67% of estimates as RCTs; the remaining 33% could not be classified reliably, although many are likely to be RCTs because Cochrane meta-analyses typically include RCT or quasi-randomised evidence.

  • measure: About 32% of estimates are standardised mean differences and 68% are differences between probit-transformed event proportions.

  • topic is Cochrane specialty and subset is source data type (published, unpublished, sought, or mixed); outcome_group is also available in the processed data.

Model of z-values

Characteristic Estimate
Probability of significance 28%
Relative probability of publication for |z| < 1.96 0.82
Successful replication for |z| > 1.96 60%
Correct sign for |z| > 1.96 97%
What do these terms mean?
Probability of significance
The reported value is the assurance: the proportion of significant results adjusted for publication bias.
Relative probability of publication
The relative probability of observing a result below the |z| = 1.96 threshold rather than above it. Values below one indicate lower observation probability below the conventional two-sided significance threshold.
Successful replication
The probability that an exact replication has the same sign and a |z| greater than 1.96, conditional on the original result having |z| greater than 1.96.
Correct sign
The probability that the observed effect has the same direction as the true effect, conditional on an original result with |z| greater than 1.96.

Cochrane Database of: Systematic Reviews mixture model plot

Data dictionary

One row is one study result within a Cochrane review analysis. Summaries count rows, not distinct reviews.

Review identifiers

Variable Definition Summary
cochrane_id Stable CD review identifier from the RM5 filename.
doi DOI of the RM5 source edition, where established; a checkpoint label is used only when no contrary evidence exists.
withdrawn Whether the source RM5 review has withdrawal status W; repeated on every study-result row. 0 733,563; 1 26,923

Review characteristics

Variable Definition Summary
specialty Cochrane Review Group code for the source edition. missing 36,034
rct Whether the matching edition abstract’s eligibility section supports RCT-only classification. TRUE 505,795; FALSE 187,878; NA 66,813
id Review identifier supplied by the RM5 parser.

Analysis identifiers

Variable Definition Summary
comparison.nr Comparison number within the review.
comparison.name Comparison label.
comparison.id Internal comparison identifier.
outcome.nr Outcome number within the comparison.
outcome.name Outcome label.
outcome.measure Effect measure named for the source outcome.
outcome.id Internal outcome identifier.
outcome.flag Source result family.
subgroup.nr Subgroup number within the outcome.
subgroup.name Subgroup label.
subgroup.id Internal subgroup identifier. missing 248,547

Study characteristics

Variable Definition Summary
study.id Internal source study identifier.
study.name Study label used as studyid in BEAR.
study.year Year parsed from the source study label or year field; implausible years are missing. missing 27,883
study.data_source Source status of the study data. PUB 630,749; MIX 101,642; SOUGHT 17,328; UNPUB 10,767

Reported effects

Variable Definition Summary
effect.size Effect estimate reported in the RM5 study table.
se Reported standard error for effect.size.
ci.lower Reported confidence interval lower limit.
ci.upper Reported confidence interval upper limit.
weight Analysis weight reported in the RM5 study table.
order Source ordering of the study result within the analysis.

Arm inputs

Variable Definition Summary
events1 Event count in arm 1 for dichotomous outcomes. missing 249,625
total1 Participant count in arm 1.
mean1 Mean in arm 1 for continuous outcomes. missing 510,861
sd1 Standard deviation in arm 1 for continuous outcomes. missing 510,861
events2 Event count in arm 2 for dichotomous outcomes. missing 249,625
total2 Participant count in arm 2.
mean2 Mean in arm 2 for continuous outcomes. missing 510,861
sd2 Standard deviation in arm 2 for continuous outcomes. missing 510,861

BEAR calculations

Variable Definition Summary
measure_group Broad effect measure category from outcome.measure.
outcome_group Heuristic outcome category from comparison, outcome and subgroup labels. efficacy 595,500; safety 100,241; bias 46,480; dropouts 18,265
phase Unused placeholder for trial phase. missing 760,486
yi Recalculated Hedges g or probit difference from arm inputs. missing 2,611
vi Sampling variance of yi. missing 2,611
measure Broad label for the recalculated effect. probit 510,861; SMD 249,625
measure_detailed Detailed label for the recalculated effect.
z yi divided by the square root of vi. missing 2,611

References

Cochrane Collaboration. 2025. Cochrane Database of Systematic Reviews. https://www.cochranelibrary.com/cdsr.
Schwab, Simon. 2024. Cochrane: Import Data from the Cochrane Database of Systematic Reviews (CDSR). https://github.com/schw4b/cochrane.