Data Citations Are No Longer Hidden References
Crossref's beta data-citation endpoint makes dataset links easier to find, but only when journals collect dataset DOIs as references or relationships instead of loose prose.
A dataset citation often arrives at a journal in the least durable place available. It is pasted into a data availability statement, mentioned in a cover letter, hidden in a supplementary file, or written as a bare repository name in the reference list. A reader may still find the data after some detective work. A machine usually cannot.
That distinction matters more after Crossref moved data citations into a dedicated beta service. STM highlighted the service on July 31, 2026, noting that the endpoint contains more than 700,000 data citations and is being updated from reference and relationship metadata deposited by Crossref members: https://stm-assoc.org/crossref-launches-new-data-citation-api-endpoint-in-beta/. Crossref is reviewing community feedback until September 2026 before deciding whether to move the endpoint into production.
The operational lesson for journals is simple and slightly uncomfortable: data citations are becoming easier to inspect from the outside. If a journal says it supports open data but its workflow leaves dataset links in prose, the public infrastructure will not rescue the record. It will expose the gap.
The Endpoint Reads Deposited Relationships
Crossref announced the data-citation endpoint in March while explaining that the Event Data public API would be sunset on April 23, 2026: https://www.crossref.org/blog/strengthening-support-for-data-citations-and-saying-goodbye-to-event-data/. The shift was not framed as a retreat from research signals. It was a narrowing of focus toward relationships between research outputs, beginning with links from scholarly works to datasets.
The documentation is useful because it states the constraint plainly: the beta endpoint returns connections between Crossref-deposited content and known datasets, using metadata found in references and relationship fields, and linking to datasets with Crossref or DataCite DOIs: https://www.crossref.org/documentation/retrieve-metadata/data-citations/. Crossref also says it does not perform matching for this service, so the metadata needs to include a dataset DOI. In the beta version, newly deposited metadata may take about five days to appear.
That makes the journal workflow decisive. A beautifully written data availability statement may help a human reader, but it will not create a data citation unless the dataset identifier is carried in a place the infrastructure can read. The same is true for a reference that names a repository record without a DOI, a supplement that includes the identifier only in a PDF, or a production workflow that strips relationship metadata before deposit.
Data Availability Is Not The Same As Data Citation
Many journals have spent the last few years improving data availability statements. That was necessary. It is not enough. A data availability statement answers where the data can be accessed, under what restrictions, or why it cannot be shared. A data citation gives the dataset a formal place in the scholarly reference chain, with creators, title, repository, version, year, and persistent identifier where available.
The 2023 joint statement from STM, DataCite, and Crossref already pointed in this direction: researchers should cite datasets in the reference section using persistent identifiers, publishers should set appropriate journal data policies, and publishers should include data citations and links to data availability statements in metadata registered with Crossref: https://www.crossref.org/blog/joint-statement-on-research-data/. The new endpoint turns that recommendation into something easier to check.
Consider two accepted articles that both comply on the surface. The first says, "Data are available at Repository X under accession Y," but never creates a structured reference or relationship. The second includes a dataset DOI in the references, preserves the data availability statement, and deposits the relationship metadata. To a reader skimming the article page, both may look responsible. To discovery systems, funders, repositories, and data stewards, they are not equivalent records.
Reference Editing Now Includes Dataset Stewardship
Data citation work often falls between departments. Editors own policy. Authors own disclosure. Production owns references. Metadata staff own DOI deposits. Repository teams own dataset records. The result is predictable: everyone supports data citation in principle, while no one owns the moment when a dataset mention becomes a citable, deposited relationship.
The immediate fix is not a broad platform rebuild. It is a narrower reference-editing rule: when an article depends on a dataset with a DOI, the dataset should be cited in a structured, reference-list form and the relationship should be preserved through metadata deposit. If the article uses a dataset without a DOI, staff should know whether the journal requires a repository that can assign one, accepts another persistent identifier, or permits a documented exception.
- Check every data availability statement for dataset identifiers, not only repository names.
- Move dataset DOIs into formal references when journal policy and discipline practice support it.
- Preserve article-to-dataset links in Crossref relationship metadata instead of leaving them only in PDF text.
- Capture dataset version, access conditions, and repository landing page before acceptance, while authors can still fix gaps.
- Document exceptions when a dataset cannot receive a DOI or cannot be publicly shared.
This is also a reviewer-service issue. Reviewers who are asked to assess reproducibility need to know whether the cited dataset is the right dataset, whether a version exists, and whether the access conditions match the claim in the manuscript. If the journal waits until proof stage to look for dataset identifiers, it has missed the moment when expert review can use them.
The Beta Is A Governance Signal
Crossref is explicit that the service is still in beta. Its documentation warns that output formats may change, users may experience downtime or slow response times, and users should include contact information through a mailto parameter while testing. Those details should keep journal teams from treating the endpoint as a settled compliance target today.
But beta status does not make the signal weak. In Crossref's May 2026 community update, the organization said the endpoint exposed more than 700,000 data citations and that, through March, Crossref typically collected 400 to 600 data citations per day: https://www.crossref.org/blog/building-refining-and-connecting-summary-of-our-may-2026-community-update/. That is enough scale for journals to ask whether their own article records are participating cleanly or relying on accidental visibility.
It is also a reminder that the scholarly record is moving from static article pages toward queryable relationships. Once data citations can be pulled by member, date range, or dataset DOI, external users can begin asking more operational questions. Which journals in a portfolio cite datasets consistently? Which repositories receive links but no formal citation? Which subject areas have data policies that do not show up in metadata? Which corrections or article updates should change a dataset relationship?
Run A Ten-Article Trace
A useful audit can be small. Pick ten research articles published in the last twelve months: three with public datasets, three with restricted-access data, two with code or software dependencies, and two where the data availability statement says data are available on request. For each article, trace the data claim across five places.
- Submission record: what did the author provide, and was a dataset DOI requested explicitly?
- Manuscript text: does the data availability statement match the reference list?
- Article outputs: do HTML, PDF, XML, and supplementary files carry the same dataset information?
- DOI metadata: can the article-to-dataset relationship be found from deposited references or relations?
- Repository record: does the dataset landing page point back to the article or at least identify the related work?
The audit will usually reveal mundane breakpoints. A DOI appears in the PDF but not in XML. A repository link redirects to a search page. A dataset version is named in the supplement but not in the article. A restricted dataset has an access committee, but the article says only that data are available on request. A copyeditor regularized the reference style but removed the identifier label that made the object recognizable.
These are small failures until a funder, institution, reader, reviewer, or research-integrity team needs to verify the evidence behind the article. Then they become delays, correspondence, corrections, or avoidable doubts about work the journal may have handled carefully in every other respect.
Practical Takeaway For Journal Leaders
Assign data-citation ownership before the next issue closes. The owner does not need to be a senior editor. The role needs authority to ask authors for missing dataset DOIs, send unclear cases back before acceptance, coordinate with production on reference treatment, and verify that article-to-dataset relationships survive in deposited metadata.
Then make one policy distinction visible to authors and staff: a data availability statement is required narrative, while a data citation is a record-level link. Journals need both when the research depends on shareable data. Crossref's beta endpoint makes the difference easier to see. Journal workflows should make the difference harder to lose.