RESEARCH PAPER DETAILS

Find out which data a research paper used and where it came from

A paper's conclusions are only as good as its data, but the data description is often split up: the source in the methods, the preprocessing in a supplement, and access terms in a data availability statement at the end. Ask Search+ "What datasets does this paper use, and can I access them?" and the answer gathers those passages with citations, so you can see the source, the time period, the filtering and the access route in the authors' own words.

Creating an account needs no payment.

Last updated October 2026

How to trace a paper's data with Search+

  1. Upload the paper with its supplement

    Data dictionaries, inclusion filters and variable lists are often moved to supplementary material, sometimes as an Excel file. Add the supplement alongside the article so one question can reach both.

  2. Ask about source, then processing, then access

    Ask where the data came from, then how it was cleaned or filtered, then how others can obtain it. Each is usually described in a different place.

  3. Follow named datasets back to their papers

    If the authors use a public dataset described elsewhere, add that dataset's own paper to the workspace and ask about collection methods across both.

Questions to ask about a paper's datasets

Source

Which datasets or databases does the paper use, and who collected them?

Time and place

What period and which regions or institutions does the data cover?

Filtering

Which records were excluded, and how many remained after cleaning?

Access

What does the data availability statement say, and is the data public, restricted or available on request?

Code

Is the analysis code available, and where do the authors say it can be found?

Across studies

Across the papers in this workspace, which ones use the same public dataset?

What to look for in a paper's data description

Primary or secondary data

Some papers collect new data; many analyse existing datasets, registries or benchmarks. The difference changes what you need to check: collection procedures for the first, version and selection for the second.

Versions and snapshots matter

Public datasets are updated over time. Look for the version, release date or access date the authors report, because results can shift between versions.

Preprocessing can change the sample

Exclusions for missing values, duplicates or outliers decide what was actually analysed. Ask how many records were removed and why.

Availability statements vary

Statements range from open repositories with identifiers, to access on reasonable request, to data that cannot be released for privacy or contractual reasons. Read the exact wording rather than assuming.

Passages, not paraphrase

Search+ finds the relevant passages by meaning, so a question about "where the data came from" can surface a methods paragraph that never uses the word dataset, with a citation you can open.

Data wording you may see, and what to ask next

Wording you may seeWhat it usually signalsA follow-up question
"publicly available at" with an identifierOpen data in a repositoryWhich version or release did the authors use?
"available from the corresponding author on reasonable request"Access controlled by the authorsDoes the paper explain what conditions apply?
"secondary analysis of"Existing data reused for a new questionWas the original data collected for a different purpose?
"after excluding"Records removed before analysisHow many were excluded at each step?
"proprietary" or "licensed from"Data others may not be able to obtainWho owns the data, and could the study be reproduced?
"code is available"Analysis scripts are publishedWhere is the code, and does it cover all reported results?

What is a dataset in a research paper?

A dataset is the collection of observations a study analyses, such as survey responses, measurements, records, images or text. A paper usually describes where it came from, what it covers, how it was processed and whether others can access it.

The dataset is not the sample size: sample size is how many units were analysed, while the dataset is the source they were drawn from. A dataset can contain far more records than the study ends up using.

Dataset questions

Can Search+ list the datasets a paper used?
Yes. Ask which data sources the paper relies on, and Search+ answers from the methods and data statements with citations, so you can check each source the authors name.
Will it find the data availability statement?
Ask whether the data and code are available. The answer cites the availability statement or the passage where access is described, including when it sits at the end of the article.
Can it tell me if the data is good enough?
It can show you what the paper reports about collection, coverage and exclusions, each with a citation. Judging whether that data supports the conclusion is your call.
Can I see which papers use the same dataset?
Yes. Upload the papers to one workspace and ask which ones use the same data source. The answer names the papers and cites the passage in each.
Can Search+ open the dataset itself?
It reads the documents in your workspace, including a data dictionary or summary table saved as an Excel file. It reports what those files say; it does not run analyses on raw data.
What if details are only in a supplement?
Upload the supplement as its own file in the same workspace, then ask again. Answers can cite both the article and the supplement.

Check the data before you trust the result

Start a workspace, upload the paper and its supplement, and ask where the data came from.

Start a workspace