Data Sharing & Publication Ethics6 min readBy Publicator Editorial

Data Sharing Ethics Need a Pre-Review Gate

COPE's September 2026 data-sharing discussion is a prompt for journals to decide consent, privacy, access, and reuse limits before reviewers inherit an unclear dataset promise.

The hardest data-sharing cases rarely look hard at submission. An author says data will be available on request. A clinical manuscript promises de-identified participant data. A qualitative study includes interview excerpts but no clear boundary around reuse. A paper using Indigenous, Tribal, community, or commercially sensitive data says access is restricted, then leaves the editor to decide whether that restriction is responsible stewardship or avoidable opacity.

COPE has put the issue directly on the September agenda. Its September 22, 2026 forum will open with a discussion on ethical considerations around data sharing, asking how researchers, editors, and institutions can realize the benefits of data sharing while meeting their ethical responsibilities: https://publicationethics.org/topic-discussions/ethical-considerations-around-data-sharing. That is a useful framing because the practical problem for journals is not simply whether data are open. It is whether the journal can explain why a particular level of openness is appropriate for this study, these participants, this consent language, and this reuse context.

The timing matters. NIH updated its Data Management and Sharing Plan page on July 22, 2026 and notes that the 2026 DMS Plan format is now required for all competing and non-competing awards. The page asks researchers to maximize appropriate sharing while accounting for legal, ethical, and technical limits, and it explicitly requires plans to describe limitations when sharing cannot be full or immediate: https://grants.nih.gov/policy-and-compliance/policy-topics/sharing-policies/dms/writing-dms-plan. Journals that wait until proof stage to think about those limits are entering the conversation too late.

The Risk Is Not Only Too Little Sharing

Journal data policies often sound as if the only failure mode is secrecy. That is understandable. Weak data availability statements, broken repository links, missing dataset identifiers, and vague "available on request" promises all damage reproducibility. But over-sharing can be an integrity problem too. Data can expose participant identities after de-identification. Consent can allow one kind of analysis and not another. Community governance can require consultation before reuse. Third-party agreements can limit redistribution. A repository can be technically open while still ethically wrong for the material.

Editors need a workflow that can hold both truths at once: data sharing is part of research transparency, and data sharing has boundaries. Treating every restriction as suspicious encourages authors to under-explain real ethical constraints. Treating every restriction as acceptable leaves readers unable to test the work. The journal's task is to make the limit inspectable, proportionate, and documented.

Make The Limits Reviewable

A data availability statement should not be the first place a journal learns that the dataset cannot be shared. By then, peer review may already have happened without enough context, reviewers may have been asked to judge claims they could not inspect, and production may have no authority to send the manuscript back for a better repository, access condition, or consent explanation.

The better control is a pre-review gate. It does not need to be heavy. It needs to separate manuscripts with routine open data from manuscripts where data access, consent, privacy, law, community governance, or commercial restrictions affect what reviewers and readers can verify. The point is not to block sensitive research. The point is to avoid discovering the sensitivity after the journal has already behaved as if the data were ordinary.

  • Ask authors whether the data include human participants, genomic information, location-sensitive records, Indigenous or community-governed material, proprietary data, or third-party licensed content.
  • Require the proposed access condition before review: open repository, controlled-access repository, mediated request, synthetic or redacted dataset, codebook only, or no sharing with stated justification.
  • Capture whether the consent, ethics approval, data use agreement, law, or repository policy supports the proposed sharing level.
  • Tell reviewers what they can inspect and what they cannot inspect, without exposing confidential material unnecessarily.
  • Record who accepted a limited-sharing rationale so the decision is not rebuilt from emails after publication.

Separate The Three Promises

Data-sharing language often compresses three different promises into one sentence. The first promise is to participants or data providers: the research team will respect consent, confidentiality, community obligations, and lawful use. The second promise is to readers: the article's claims can be evaluated with enough evidence to judge the work. The third promise is to data creators: reuse will preserve credit, attribution, and context rather than turning data into unattributed raw material.

Those promises can conflict. A clinical dataset may be scientifically valuable and still unsuitable for open release. A community dataset may be shareable only under governance terms that a generic repository cannot express. A proprietary dataset may support an article if reviewers can inspect the analysis route, but readers may need a clear statement of what cannot be checked independently. A journal that forces all of those cases into the same template will produce statements that are neat, misleading, or both.

The editorial file should preserve the distinction. What did the participants or data providers permit? What can reviewers verify? What will readers receive? What reuse is allowed? What credit or citation is required? If a correction, expression of concern, or reader challenge arrives later, these are the questions the journal will need to answer.

Clinical Trial Statements Set A Useful Bar

ICMJE's clinical-trial recommendations are a good model even outside clinical medicine. For manuscripts reporting clinical trials, ICMJE requires a data-sharing statement and says it should state whether de-identified individual participant data will be shared, what data will be shared, whether related documents will be available, when and for how long data will be available, and by what access criteria and mechanism: https://www.icmje.org/recommendations/browse/publishing-and-editorial-issues/clinical-trial-registration.html.

That structure matters because it replaces virtue language with operational commitments. "Data available on reasonable request" is not enough if nobody defines reasonable, identifies the data, names the decision maker, sets a time window, or explains the access route. It also matters when a data-sharing plan changes. ICMJE says changes after registration should be reflected in the submitted and published statement and updated in the registry record. Journals in other fields can borrow the habit: if the access promise changes, the public record should change with it.

Data Citation Is An Ethics Tool Too

The FORCE11 Joint Declaration of Data Citation Principles is often invoked for discovery and credit, but it also helps with ethics. The principles treat data as legitimate citable research products, emphasize credit and attribution, call for access to data and relevant metadata or documentation, and stress specificity, verifiability, provenance, and fixity: https://force11.org/info/joint-declaration-of-data-citation-principles-final/. Those ideas are not separate from ethical sharing. They are how a journal prevents reuse from stripping away context.

A cited dataset with a persistent identifier, version, access condition, creators, repository landing page, and reuse terms is easier to govern than a nameless file mentioned in prose. It gives credit to data generators. It lets readers see whether they are looking at the same material the authors used. It gives repositories a role in access control. It also lets journals document cases where the metadata can be public even when the data themselves require mediation or cannot be released.

Build The Gate Before Peer Review

The operational change is small but consequential: data ethics should be triaged before reviewers are invited. A managing editor can ask a short set of questions at intake, route sensitive cases to an editor or data specialist, and decide whether reviewers need access to a repository, a redacted file, a codebook, a synthetic dataset, an analysis script, or a stronger limitation statement. That is much easier before review than after acceptance.

This also protects reviewers. They should not be asked to download sensitive participant material through ad hoc links, negotiate access terms directly with authors, or infer from a vague statement whether the data can support the paper. The journal should define the review condition: what is available, under what confidentiality terms, what cannot be shared, and what the reviewer is expected to evaluate anyway.

Practical Takeaway for Journal Leaders

Run a one-month audit of data-sharing decisions before the COPE forum. Pull ten recent submissions with data availability statements, including at least three that limit access. For each, check whether the journal captured the data type, consent or legal basis, repository or access mechanism, reviewer access condition, public statement, dataset citation or identifier, and the person who approved any limitation.

If the answer is scattered across cover letters, reviewer notes, production queries, and private emails, the journal does not yet have data-sharing governance. It has data-sharing prose. Ethical sharing needs a gate early enough to shape peer review, precise enough to protect participants and data creators, and visible enough that readers can understand what evidence is available and why some evidence may responsibly remain controlled.