返回 Skills 目录
yuan1z0825/nature-skills已通过检查

SKILL DETAIL

nature-data

yuan1z0825/nature-skills/nature-data

该技能用于帮助作者准备、审计或修订符合Nature期刊要求的数据可用性声明、数据存储库计划、数据集引用和FAIR元数据清单。它适用于用户询问Nature数据可用性、研究数据共享、存储库选择、登录号、受限或敏感数据、源数据、补充数据集、DataCite风格的数据集引用、学术出版的FAIR元数据,或中文作者准备Nature系列投稿时的中英文数据可用性措辞。 该技能还适用于一般学术写作中的数据需求,即使不涉及“Nature”,例如为任何期刊撰写数据可用性声明、代码/数据共享部分、论文写作中的存储库选择,以及中文表述如“数据可用性声明”、“数据可用性”、“数据共享”、“代码可用性”、“学术写作数据声明”、“写数据声明”、“数据存放”、“数据仓库选择”。

安装量 · 145查看来源

Installation

npx skills add https://github.com/yuan1z0825/nature-skills --skill nature-data

技能文件

SKILL.md

最近同步 · 2026年8月27日

agents/openai.yaml
interface:
  display_name: "Nature Data"
  short_description: "Draft bilingual-aware Nature data statements"
  default_prompt: "Use $nature-data to draft a Nature-ready Data Availability statement, repository plan, and FAIR checklist."
manifest.yaml
name: nature-data
version: 2.2.0
description: >
  Declarative manifest for the static/dynamic split. SKILL.md uses this to
  decide which fragments to load for a data-availability request.

# Design note: nature-data is a linear, parameterised workflow (inventory ->
# classify access route -> choose repository/identifier -> draft statement ->
# cite datasets -> FAIR audit). Its variation — journal, access route, user
# language — is handled at runtime by inline rules and the on-demand references,
# not by loading different large content bodies. There is therefore no content
# axis; the split is core (always loaded) plus on-demand references. nature-data
# does not use the prose-oriented nature-shared layer (it governs data-availability
# policy, with its own source basis).

always_load:
  - static/core/stance.md
  - static/core/chinese-mode.md
  - static/core/workflow.md

references:
  on_demand:
    - condition: governing Nature/Springer Nature data-sharing rules or edge-case policy logic
      path: references/policy-principles.md
    - condition: target is the flagship journal Nature, or exact Nature Article data, code, materials, mandatory-deposition, reviewer-access, or statement-placement rules are needed
      path: references/nature-article-requirements.md
    - condition: target is Nature Machine Intelligence, or exact NMI Data Availability, Code availability, repository, reviewer-access, central-code, software-checklist, or statement-placement rules are needed
      path: ../nature-shared/journal-formats/nature-machine-intelligence.md
    - condition: the user writes in Chinese, needs bilingual wording, or provides Chinese availability notes
      path: references/chinese-author-alignment.md
    - condition: ready-to-adapt Data Availability statement patterns
      path: references/statement-patterns.md
    - condition: repository choice, accession, DOI, embargo, versioning, or dataset citation guidance
      path: references/repository-and-identifiers.md
    - condition: FAIR checks, README metadata, file organization, licences, provenance, or DataCite fields
      path: references/fair-metadata-checklist.md
    - condition: justify rules with official sources or check which source supports which rule
      path: references/source-basis.md
README_EN.md
# `nature-data` Skill

[中文说明](README.md)

`nature-data` prepares, audits, or revises Nature / Springer Nature-style Data Availability statements, repository plans, dataset citations, and FAIR metadata checklists.

## What To Use It For

- Draft a manuscript-ready Data Availability statement.
- Check whether statements such as "data are available from the corresponding author" or "raw data are in supplementary materials" are sufficient.
- Map every dataset supporting a conclusion to a repository, accession, DOI, license, or access condition.
- Distinguish public data, controlled-access data, third-party data, supplementary-data cases, and not-applicable cases.
- Turn Chinese author notes into submission-ready English and list facts that still need confirmation.
- Check flagship `Nature Article` statement placement, mandatory repositories, reviewer access, central code, materials, and structure-data files.
- Check `Nature Machine Intelligence` (NMI) Data Availability, the following separate Code availability section, review-stage data/central-code access, restrictions, and the Software Submission Checklist.

## Typical Requests

- "Write this manuscript's Data Availability statement in Nature style."
- "Some data cannot be public; draft a controlled-access statement."
- "Check whether my data statement is missing accessions, repositories, or licenses."

## What You Need To Provide

- The data source supporting each figure, table, or conclusion.
- Upload status, repository link, accession, DOI, embargo, or access restriction.
- Availability boundaries for code, materials, protocols, and third-party data.

## Outputs

- Ready-to-paste English Data Availability statement.
- Dataset-to-figure/conclusion mapping table.
- Missing-information checklist and FAIR / DataCite metadata check.
- Conservative wording for restricted, third-party, or supplementary-data cases.

## Boundaries

- The skill does not invent accessions, DOIs, licenses, repository records, or access restrictions.
- Missing information is handled with a usable draft plus a short confirmation checklist.
- Data restricted by ethics, privacy, commercial terms, or third-party agreements need real restriction details from the author.

## Related Skills

- `nature-experiment-log`: organize experiment logs and raw attachments into traceable data sources.
- `nature-statistics`: check statistical reporting and source-data wording.
- `nature-response`: respond to reviewer concerns about data availability.
README.md
# `nature-data` 技能

[English](README_EN.md)

`nature-data` 用于准备、审查或修改 Nature / Springer Nature 风格的 Data Availability statement、数据仓储计划、数据集引用和 FAIR 元数据清单。

## 适合用它做什么

- 起草可放入稿件的 Data Availability statement。
- 审查“数据可向通讯作者索取”“原始数据见补充材料”等表述是否充分。
- 为每个支撑论文结论的数据集匹配仓储、accession、DOI、许可证或访问条件。
- 区分公开数据、受控访问数据、第三方数据、补充材料数据和不适用场景。
- 将中文作者笔记转成投稿可用英文,并列出需要作者确认的信息。
- 对旗舰 `Nature Article` 检查声明位置、强制仓储、审稿访问、中心代码、材料和结构数据文件。
- 对 `Nature Machine Intelligence`(NMI)检查 Data Availability、其后独立的 Code availability、审稿期数据/中心代码访问、限制条件和 Software Submission Checklist。

## 典型请求

- “帮我把这篇稿子的 Data Availability 写成 Nature 风格。”
- “这些数据有一部分不能公开,帮我写受控访问说明。”
- “检查我的数据声明是否缺 accession、仓储或许可证。”

## 你需要提供

- 支撑每个图表或结论的数据来源。
- 数据是否已上传、仓储链接、accession、DOI、embargo 或访问限制。
- 代码、材料、实验方案和第三方数据的可用性边界。

## 产出

- 可粘贴的英文 Data Availability statement。
- 数据集到图表/结论的映射表。
- 缺失信息清单和 FAIR / DataCite 元数据检查。
- 对受限数据、第三方数据或补充材料数据的保守表述建议。

## 边界

- 不会编造 accession、DOI、许可证、仓储记录或访问限制。
- 信息缺失时,会给出可用草稿和短确认清单,而不是假装完整。
- 受伦理、隐私、商业或第三方协议限制的数据,需要作者提供真实限制条件。

## 相关技能

- `nature-experiment-log`:把实验记录和原始附件整理成可追溯数据来源。
- `nature-statistics`:检查统计报告和 source data 表述。
- `nature-response`:回应审稿人关于数据可用性的质疑。
references/chinese-author-alignment.md
# Chinese Author Alignment

Use this file when the user writes in Chinese, provides a Chinese Data Availability draft, or asks
for bilingual wording. The goal is not to translate Chinese literally. The goal is to convert the
author's Chinese description into a Nature-ready English availability route.

## Core terminology

| 中文 | Preferred English | Notes |
|---|---|---|
| 数据可用性声明 / 数据获取声明 | Data Availability | Use the journal heading `Data Availability`. |
| 本研究产生的数据 | data generated in this study | Include repository and identifier when public. |
| 原始数据 | raw data | Do not call processed tables raw data. |
| 处理后数据 | processed data | State whether processing scripts are available. |
| 源数据 | source data | Usually data underlying figures or tables. |
| 补充材料 / 附录 | Supplementary Information | Use exact file/table names when possible. |
| 公共数据库 | public database / public repository | Name the database and identifier. |
| 数据存储库 | data repository | Prefer repository over platform unless it is a true archive. |
| 登录号 / 编号 | accession number | Use for repositories that assign accession IDs. |
| DOI / 永久链接 | DOI / persistent URL | Prefer DOI when available. |
| 受限数据 | restricted data | Explain legal, ethical, consent, commercial, or third-party reason. |
| 脱敏数据 | de-identified data | Do not say anonymous unless re-identification risk is addressed. |
| 合理请求 | reasonable request | Not enough alone; add route, eligibility, and conditions. |
| 通讯作者 | corresponding author | Avoid making an email the only durable access route if an institutional route exists. |
| 数据使用协议 | data-use agreement | State when required for access. |
| 伦理审批 | ethics approval | Name approval body or requirement when relevant. |
| 代码可用性 | Code Availability | Keep separate if the journal separates data and code. |

## Chinese-to-English conversion rules

- Convert "本文所有数据均包含在正文和补充材料中" to a specific claim:
  name Source Data files, Supplementary Tables, or repository records. If raw data are absent, say
  so as a risk flag rather than pretending they are included.
- Convert "可向通讯作者合理索取" only after adding:
  why public sharing is impossible, who reviews requests, eligible requesters, required approvals
  or data-use agreement, and expected access route.
- Convert "数据因隐私原因不可公开" into a controlled-access pattern:
  state privacy/consent/legal basis, public metadata if available, access committee or institution,
  and conditions.
- Convert "商业数据/企业数据不可公开" into a third-party or commercial restriction pattern:
  name the provider or owner, request route, and whether derived or aggregate data can be shared.
- Convert "数据将在接收后上传" into an action item:
  deposit before submission or create a private reviewer link if the repository supports it.
- Convert "使用公开数据集" into a citation requirement:
  include source, version/release/date accessed when relevant, and dataset citation.

## Bilingual intake questions

Ask only what is needed for the statement.

```text
请确认这些字段:
1. 哪些数据支撑主文图、补充图和统计分析?
2. 每类数据是否已有仓库、DOI、登录号或审稿人私密链接?
3. 是否包含人类参与者、隐私、商业、第三方授权或国家/机构限制?
4. 如果数据不能公开,谁负责审核申请?需要伦理审批或数据使用协议吗?
5. 是否有代码、脚本或 README 能解释 raw data 到 figure source data 的处理过程?
```

## Common Chinese draft fixes

| 中文原意 | Avoid literal English | Nature-ready direction |
|---|---|---|
| 数据可向通讯作者索取。 | Data are available from the corresponding author upon request. | State the restriction reason and institutional access process. |
| 所有数据见补充材料。 | All data are in the supplementary materials. | Name exact Supplementary Tables/Source Data and flag missing raw data if any. |
| 数据暂未上传。 | Data will be uploaded later. | Deposit now or list repository action as blocking. |
| 使用了公开数据库。 | Public databases were used. | Name database, accession/version/date accessed, and cite dataset. |
| 因隐私不能公开。 | Data cannot be public for privacy reasons. | Add de-identification status, access committee, eligibility, and agreement terms. |

## Recommended bilingual output

When useful, provide English first and Chinese second:

```text
Data Availability
[English statement for submission]

中文核对
- 这句话对应中文含义:[brief Chinese explanation]
- 需要作者确认:[missing accession / repository / ethics condition]
```

Do not put Chinese explanatory notes inside the final English statement unless the target journal
allows bilingual manuscript text.
references/fair-metadata-checklist.md
# FAIR Metadata Checklist

## Contents

- [Quick FAIR test](#quick-fair-test)
- [DataCite core fields](#datacite-core-fields)
- [Dataset README template](#dataset-readme-template)
- [Summary](#summary)
- [Files](#files)
- [Variables and units](#variables-and-units)
- [Methods and provenance](#methods-and-provenance)
- [Software and environment](#software-and-environment)
- [Access and licence](#access-and-licence)
- [Citation](#citation)
- [File organization](#file-organization)
- [Provenance prompts](#provenance-prompts)
- [Licence guidance](#licence-guidance)
- [Final audit](#final-audit)


Use this file to audit whether a dataset deposit is findable, accessible, interoperable, and
reusable enough for a Nature-style submission.

## Quick FAIR test

| Principle | Practical check |
|---|---|
| Findable | Dataset has a persistent identifier, rich title/abstract/keywords, searchable repository record, and metadata that names the data identifier. |
| Accessible | Identifier resolves through a standard protocol; access conditions are explicit; metadata stay public even if data are restricted. |
| Interoperable | Files use community formats where possible; metadata use shared vocabulary, units, identifiers, and qualified links to related data/code/publication. |
| Reusable | Licence, provenance, methods, variables, quality-control notes, version, and community-standard metadata are clear enough for reuse. |

## DataCite core fields

Mandatory fields commonly expected for DOI-style dataset records:

- Identifier
- Creator
- Title
- Publisher / repository
- Publication year
- Resource type

Strongly recommended when available:

- contributor and role
- description / abstract
- subject keywords
- funding reference
- related identifiers: manuscript preprint/article, code repository, protocol, previous dataset
- version
- licence / rights
- geolocation or temporal coverage for spatial/temporal data
- language

## Dataset README template

```text
# [Dataset title]

## Summary
[One-paragraph description of what the dataset contains and which manuscript results it supports.]

## Files
- [filename]: [contents, format, size, related figure/table]

## Variables and units
[Column/field name] | [definition] | [unit] | [allowed values/missing-value code]

## Methods and provenance
[How data were generated, collected, transformed, filtered, normalised, or aggregated.]

## Software and environment
[Software, package versions, scripts, notebooks, operating system or instrument software when relevant.]

## Access and licence
[Licence, access restrictions, data-use agreement, embargo, or controlled-access process.]

## Citation
[Preferred dataset citation.]
```

## File organization

- Use stable, descriptive filenames instead of local shorthand.
- Keep raw and processed data separate.
- Include a manifest for archives or large multi-file deposits.
- Map source data to exact figure panels and table numbers.
- Preserve units in column names or data dictionaries, not only in manuscript captions.
- Record missing-value codes and filtering decisions.
- Include checksums for large or critical files when the repository does not generate them.

## Provenance prompts

Ask the author:

- What instrument, survey, simulation, database, or processing pipeline produced each file?
- Which script or notebook converts raw data into each figure or statistical table?
- Which samples, time points, conditions, or participants were excluded, and why?
- What version of each third-party dataset was used?
- Are there licences, consent forms, data-use agreements, or ethics approvals that limit reuse?
- Has any data been transformed in a way that prevents reconstruction of the raw values?

## Licence guidance

- Prefer a standard open licence when data can be public.
- Use the repository's licence field rather than only writing licence text in the manuscript.
- Use CC0 or CC-BY-style terms only when appropriate for the data and institution.
- Do not apply an open licence to third-party or participant data unless the authors hold the right
  to do so.
- For code, use a software licence and archive a release when possible.

## Final audit

Block submission until these are resolved:

- no Data Availability statement for original research
- no identifier or stable access route for data supporting central conclusions
- sensitive data restriction without access procedure
- third-party data with no source or permission route
- public dataset with no licence or README
- claim that data are in the paper when figure source data are absent
- mismatch between manuscript statement, repository record, and supplementary files
references/nature-article-requirements.md
# Flagship Nature Article data, code and materials requirements

Use this reference when the target is the flagship journal **Nature**. It adds
journal-specific submission placement and specialist checks to the general
repository and FAIR workflow.

## Contents

1. Data Availability
2. Mandatory and specialist deposition gate
3. Code Availability
4. Materials and protocols
5. Controlled, clinical and third-party data
6. Structure-data submission files
7. Nature readiness output
8. Official sources

## Data Availability

- Every original-research Article needs a headed `Data Availability` statement.
- Place it with the Methods/end-of-Methods material according to the current
  Nature Article sequence.
- Map primary and reused datasets to repositories, accession numbers or exact
  Source Data/Supplementary files.
- Make the minimum dataset needed to interpret, verify and extend the work
  transparent.
- Supply supporting data to editors and referees when requested.
- Require full public access at publication unless a disclosed, justified
  restriction applies.
- Formally cite deposited datasets in the reference list using author, title,
  repository and full identifier URL.

Large datasets should normally use repositories rather than Supplementary
Information. A personal site, mutable cloud folder or unarchived GitHub branch
is not a durable repository record.

## Mandatory and specialist deposition gate

Check community repositories and accession numbers for applicable data,
including:

- protein sequences
- DNA/RNA sequences and sequencing data
- genetic polymorphisms and linked genotype/phenotype data
- macromolecular structures and electron-microscopy maps
- gene-expression data, including MIAME compliance where applicable
- small-molecule crystallography
- proteomics
- Earth, space and environmental data where a community repository exists

This list is a routing reminder, not a substitute for the current mandatory
repository table. Verify the current official page before naming the required
repository.

## Code Availability

For previously unreported custom code, algorithms or software central to the
main claims:

- make the code available to editors and referees on request during assessment
- include a separate headed `Code Availability` statement
- state how the code can be accessed and every restriction
- prefer an archived, versioned, DOI-minting repository for publication
- cite the archived software record in the reference list
- do not claim readiness when a central code dependency is unavailable and the
  editor has not accepted the restriction

Keep code version, environment, licence, model weights and execution data
separate in the inventory when they have different access routes.

## Materials and protocols

- Disclose any restriction on unique materials at submission and in the
  manuscript.
- Name who will handle material requests when it is not the corresponding
  author.
- Use RRIDs or other persistent identifiers for key antibodies, cell lines,
  organisms and tools where available.
- Report cell-line source, authentication and distribution restrictions.
- Deposit step-by-step protocols on a citable platform when available and cite
  the DOI or stable record in Methods.

## Controlled, clinical and third-party data

For controlled access, state:

- why access is restricted
- the responsible committee or institutional route
- eligibility and application procedure
- expected response timeframe
- data-use-agreement and reuse restrictions
- what discoverable metadata remain public

For clinical trial data, also state what de-identified participant data and
documents will be shared, when and for how long, with whom, for which analyses,
and by what mechanism. `Undecided` is not an acceptable final sharing plan.

For third-party or proprietary data, identify the provider to editors, disclose
the restriction and confirm that peer-review and post-publication verification
access is permitted under the stated terms.

## Structure-data submission files

When small-molecule crystallography applies, coordinate with the shared
research-compliance checklist for `.cif`, structure factors, probability-
ellipsoid artwork and CheckCIF output. For macromolecular structures, confirm
the required validation report and release-on-publication status.

## Nature readiness output

Add these columns to the standard data audit:

| Dataset/code/material | Claim supported | Required repository/file | Reviewer access | Publication access | Statement placement | Status |
|---|---|---|---|---|---|---|

Use `blocked` for missing mandatory deposition, inaccessible claim-critical
data/code, undefined controlled access, or missing structure validation files.

## Official sources

Verified 2026-08-08:

- Nature initial submission: <https://www.nature.com/nature/for-authors/initial-submission>
- Nature Portfolio reporting standards and availability: <https://www.nature.com/nature-portfolio/editorial-policies/reporting-standards>
- Nature forms and declarations: <https://www.nature.com/nature/for-authors/forms-and-declarations>
references/policy-principles.md
# Policy Principles

## Contents

- [Governing rules](#governing-rules)
- [Minimal dataset test](#minimal-dataset-test)
- [Availability routes](#availability-routes)
- [Data, code, materials, protocols](#data-code-materials-protocols)
- [Sensitive and human-participant data](#sensitive-and-human-participant-data)
- [Submission-stage checks](#submission-stage-checks)
- [Source notes](#source-notes)


Use this file when deciding what a Nature-ready data statement must disclose.

## Governing rules

- Every original research article needs a Data Availability statement.
- The statement must say what supporting data exist, where they can be found, and any access
  conditions.
- The statement must cover data generated by the study and secondary data reused for analysis.
- Public repository deposition is preferred. For community-mandated data types, use the required
  repository.
- Reviewers may need access to underlying data and code during evaluation.
- Restrictions are allowed only when they are justified and disclosed. Privacy, consent, endangered
  locations, third-party licences, commercial restrictions, and national law are common reasons.
- Restricted data still need a durable access route: named data access committee, institution,
  controlled-access repository, application procedure, or responsible group.
- The statement should not hide key evidence in vague language such as "data available upon
  reasonable request" unless the reason and process are explicit.

## Minimal dataset test

Ask whether an independent reader can inspect or reproduce the paper's central findings from the
available material.

Include:

- source data for main figures and key supplementary figures
- raw or sufficiently reusable data, according to community norms
- processed data used for statistics, plots, model training, or validation
- analysis-ready tables if raw data require specialized transformation
- third-party datasets with source, version, date accessed when relevant, and licence/access terms
- representative metadata for restricted datasets, even when records themselves cannot be public

Exclude only when defensible:

- data that were not used to support a result
- purely theoretical work that generated or analysed no dataset
- identifiable human data that cannot be anonymised or shared under consent and law

## Availability routes

Use one route per dataset or dataset family.

| Route | Use when | Statement must include |
|---|---|---|
| Public repository | Data can be openly shared | repository, DOI/accession, dataset title or scope, licence if known |
| Controlled repository | Data are sensitive but discoverable | repository, accession/record, access committee or procedure, restrictions |
| Supplementary/source data | Small supporting files are hosted with paper | exact file/table/source-data mapping |
| Reused public data | The study analyses existing public data | original repository/source, identifier, version/date accessed if needed |
| Third-party restricted | Data are licensed or owned by another party | owner/source, why not public, request route, permission condition |
| Request-based access | No repository route is possible | reason, responsible group, eligibility, expected conditions, contact route |
| Not applicable | No datasets were generated or analysed | concise reason; do not use for studies with any empirical data |

## Data, code, materials, protocols

Data Availability is not a substitute for code, materials, or protocol availability.

- Put custom code in a Code Availability section when the journal separates it.
- Mention code in Data Availability only when it is bundled with the dataset and needed to interpret
  files.
- For unique biological materials, reagents, cell lines, plasmids, or model organisms, use
  persistent identifiers where available and state distribution restrictions separately.
- For protocols, cite protocol repositories or include enough method detail for reproducibility.

For a flagship Nature Article, load `nature-article-requirements.md` and apply
its separate Data Availability/Code Availability placement, mandatory-deposition
and reviewer-access gates.

## Sensitive and human-participant data

For sensitive data, preserve transparency without breaching consent or law.

State:

- why open sharing is not possible
- whether anonymised, aggregate, synthetic, or representative data can be shared
- where metadata or a summary record is available
- who reviews access requests
- what approval, data-use agreement, or ethics condition applies
- whether access is limited to non-commercial, academic, local-jurisdiction, or qualified users

Avoid:

- naming a single individual as the only durable access route when an institutional route exists
- implying data are available if access depends on impossible or undefined permissions
- promising public release later without a repository, date, and responsible party

## Submission-stage checks

Before finalizing, confirm:

- all accession numbers, DOIs, and URLs resolve
- embargoed/private reviewer links work anonymously where required
- restricted data metadata records are public if the records themselves are not
- supplementary files match statement wording
- data citations appear in the reference list where the journal expects them
- no claim depends on unavailable data without explanation

## Source notes

- Springer Nature research data policy requires Data Availability statements for original articles
  and asks authors to describe available data, location, and access terms.
- Nature Portfolio reporting standards require prompt availability of data, materials, code, and
  associated protocols, with restrictions disclosed to editors at submission.
- Scientific Data policy favours repository deposition, especially for primary data, and requires
  repository hosting for Data Descriptor datasets.
references/repository-and-identifiers.md
# Repository and Identifiers

Use this file when selecting repositories, checking accession strategy, or writing dataset
citations.

## Repository decision tree

1. Use a mandated repository when the data type requires it.
2. If no mandate applies, use a discipline-specific, community-recognised repository.
3. If no domain repository fits, use a trusted generalist or institutional repository that provides
   persistent identifiers and durable metadata.
4. Do not use personal websites, lab websites, ad hoc cloud folders, or unpublished private drives as
   the only availability route.
5. For very large data, use a repository or institutional infrastructure that can preserve metadata
   and provide clear access instructions even if bulk files require special transfer.

## What a repository record should provide

- persistent identifier: DOI, accession, Handle, ARK, or equivalent stable record
- public landing page with title, creators, abstract/description, repository, date, version, licence
- file list with sizes and formats
- README or data dictionary
- provenance and processing description
- relation to the manuscript and related code
- clear access procedure for restricted data
- versioning or update policy

## Common repository categories

Choose according to field norms; this list is not exhaustive.

| Data type | Typical repository pattern |
|---|---|
| Sequencing / gene expression | GEO, SRA, ENA, ArrayExpress or field-specific omics archive |
| Protein/nucleic acid structures | wwPDB / PDB |
| Small-molecule crystallography | CCDC or other crystallographic archive required by the journal |
| Proteomics | PRIDE or ProteomeXchange member repository |
| Metabolomics | MetaboLights or domain archive |
| Neuroimaging | OpenNeuro, DANDI, NDA, or controlled-access archive when required |
| Clinical or sensitive human data | controlled-access repository such as dbGaP, EGA, controlled institutional archive, or data access committee |
| Earth/environment/space science | PANGAEA, NASA/NOAA/ESA data centres, domain observatories |
| Social science | ICPSR, Dataverse, UK Data Service, OpenICPSR, OSF where appropriate |
| General datasets | Dryad, Zenodo, Figshare, OSF, institutional repository with DOI support |

Always check the target journal and funder because some data types have mandatory repositories.

## Identifier rules

- Prefer final public identifiers before submission.
- If the record is private during review, provide an anonymous reviewer link when the repository
  supports it.
- Do not cite temporary sharing links as dataset identifiers.
- Include accession numbers exactly as assigned by the repository.
- Use one identifier per coherent dataset record; avoid burying unrelated data under one unclear DOI.
- Version datasets when files change after review or publication.
- If the dataset has a DOI, cite the DOI rather than only the repository URL.

## Dataset citation pattern

Dataset references should include the minimum DataCite-style elements:

```text
[Creator(s)] ([Publication year]) [Dataset title]. [Repository]. [Identifier].
```

Add version when meaningful:

```text
[Creator(s)] ([Year]) [Dataset title], version [version]. [Repository]. [DOI/accession].
```

For reused public data, cite the dataset in the reference list when the dataset supports conclusions.
Mentioning it only in the Data Availability statement may be insufficient.

## Repository readiness checklist

Before submission:

- DOI/accession resolves to the intended landing page
- title matches manuscript terminology
- creators and affiliations are correct
- licence is present and compatible with intended reuse
- files open without proprietary software where possible
- README explains columns, units, missing values, transformations, and scripts
- figure source data are clearly mapped to figure panels
- restrictions and access conditions match the manuscript statement
- embargo/private links have been tested outside the author account

## Red flags

- "Data available on GitHub" without release DOI or archive
- repository record has no licence
- uploaded zip file has no README or file manifest
- accession exists but is not public, not under embargo, and not available to reviewers
- filenames use local analysis shorthand that readers cannot interpret
- manuscript cites one dataset but results depend on several unlisted secondary sources
references/source-basis.md
# Source Basis

Use this file when a user asks why a rule exists, wants primary-source justification, or needs to
audit the `nature-data` skill against real policy sources.

## Source map

| Skill rule | Primary support |
|---|---|
| Original research needs a Data Availability statement. | Springer Nature research data policy says original articles must include a data availability statement and that it should describe available data, location, and access terms. |
| The statement must cover original and reused data, including data that cannot be public. | Springer Nature policy applies to datasets needed to interpret and replicate conclusions and explicitly includes original/reused data and non-publicly shareable data. |
| Supporting data should be public where possible, with mandatory community repositories for some data types. | Springer Nature policy strongly encourages public availability for datasets supporting analysis and conclusions and mandates sharing for community-endorsed data types. |
| Reviewers may need access to underlying data and code. | Springer Nature policy states peer reviewers are entitled to request access to underlying data and code when needed for evaluation. |
| Nature-style statements must expose the minimum dataset needed to interpret, verify, and extend the work. | Nature Portfolio reporting standards describe transparent access conditions for the minimum dataset needed to interpret, verify, and extend research. |
| Materials, data, code, and protocols should be available without undue qualifications, and restrictions must be disclosed. | Nature Portfolio reporting standards state availability is a publication condition and restrictions must be disclosed at submission and in the manuscript. |
| Repositories are preferred over large supplementary files. | Nature Portfolio reporting standards discourage large datasets in supplementary information and prefer repositories; Scientific Data also strongly encourages repository deposition, especially for primary data. |
| Repository choice should prefer discipline-specific, community-recognised repositories, with generalist or institutional repositories as fallback. | Springer Nature repository guidance recommends discipline-specific community repositories where possible, otherwise generalist or institutional repositories. |
| Sensitive data should use safe sharing, controlled access, metadata records, or trusted environments where appropriate. | Springer Nature sensitive data guidance recommends repository use where possible, controlled-access repositories, trusted research environments, and metadata records for non-public data. |
| Human, non-human sensitive, proprietary, and third-party data need explicit rights and access logic. | Springer Nature sensitive data guidance lists identifiable human data, other sensitive data, and proprietary/third-party data as categories requiring special handling. |
| Rawness and reusability should follow community norms. | Scientific Data policy says data should be provided at a level of rawness allowing reuse in line with accepted community norms. |
| FAIR checks should include findability, accessibility, interoperability, and reusability for humans and machines. | Wilkinson et al. formally describe the FAIR principles and emphasize findable, accessible, interoperable, reusable digital objects for people and machines. |
| Dataset citation metadata should include persistent identifiers and core descriptive fields. | DataCite Metadata Schema defines core metadata properties for accurate and consistent identification, citation, and retrieval of resources. |

## Official sources

- Springer Nature, Research data policy:
  <https://www.springernature.com/gp/journal-policies/15369670>
- Springer Nature, Data availability statements:
  <https://www.springernature.com/gp/authors/research-data-policy/data-availability-statements>
- Springer Nature, Data repository guidance:
  <https://www.springernature.com/gp/authors/research-data-policy/recommended-repositories>
- Springer Nature, Sensitive data:
  <https://www.springernature.com/gp/authors/research-data-policy/sensitive-data>
- Nature Portfolio, Reporting standards and availability of data, materials, code and protocols:
  <https://www.nature.com/nature-portfolio/editorial-policies/reporting-standards>
- Example Nature Portfolio journal reporting standards page:
  <https://www.nature.com/npj2dmaterials/editorial-policies/reporting-standards>
- Nature Research, Data availability statements and data citations policy FAQ:
  <https://www.nature.com/documents/nr-data-availability-statements-data-citations-faqs.pdf>
- Scientific Data, Data policies:
  <https://www.nature.com/sdata/policies/data-policies>
- Wilkinson et al. 2016, The FAIR Guiding Principles for scientific data management and stewardship:
  <https://www.nature.com/articles/sdata201618>
- DataCite Metadata Schema:
  <https://schema.datacite.org/>

## Notes for future updates

- Check target journal instructions first because Nature Portfolio journals can add field-specific
  requirements.
- Check DataCite's latest schema before naming version-specific fields. As of 2026-05-01, the
  DataCite schema landing page lists Metadata Schema 4.7 as the latest release.
- Keep this file as a source map, not a long policy mirror. Link to official pages rather than
  copying full policy text.
references/statement-patterns.md
# Statement Patterns

## Contents

- [Public repository, single dataset](#public-repository-single-dataset)
- [Public repository, multiple datasets](#public-repository-multiple-datasets)
- [Data in paper and supplementary files only](#data-in-paper-and-supplementary-files-only)
- [Reused public data](#reused-public-data)
- [Mixed generated and reused data](#mixed-generated-and-reused-data)
- [Controlled-access human or sensitive data](#controlled-access-human-or-sensitive-data)
- [Third-party or licensed data](#third-party-or-licensed-data)
- [Commercially restricted data](#commercially-restricted-data)
- [Embargoed data](#embargoed-data)
- [Request-based access with justified restriction](#request-based-access-with-justified-restriction)
- [No datasets generated or analysed](#no-datasets-generated-or-analysed)
- [Anti-patterns to revise](#anti-patterns-to-revise)
- [Audit questions](#audit-questions)


Use these patterns as starting points. Replace bracketed fields with verified information. Delete
any sentence that does not apply.

For Chinese users, treat the Chinese line under each pattern as author-facing guidance, not as
submission text. Submit the English statement unless the journal explicitly asks otherwise.

## Public repository, single dataset

```text
The [raw/processed/source] data supporting the findings of this study are available in
[Repository] under accession [ACCESSION] / at [DOI or persistent URL]. The deposited record
contains [brief contents: e.g. raw measurements, processed tables, figure source data, metadata
and analysis inputs].
```

中文对应:本研究的原始/处理后/源数据已存储在某个正式仓库,并有登录号、DOI 或永久链接。

## Public repository, multiple datasets

```text
The datasets generated in this study are available as follows: [dataset family 1] in
[Repository] under [DOI/accession]; [dataset family 2] in [Repository] under [DOI/accession];
and figure source data in [Repository/Supplementary Data file] under [identifier or file name].
```

中文对应:不同类型数据分别放在不同仓库或文件中,需要逐一说明,不能笼统写“数据见附件”。

## Data in paper and supplementary files only

Use only when the supporting dataset is genuinely small and fully represented in the article,
source data, or supplementary files.

```text
All data supporting the findings of this study are included in the paper, its Supplementary
Information, and Source Data files. [Name exact Supplementary Tables/Data files when possible.]
```

中文对应:只有当支撑结论的数据确实都在正文、补充材料和 Source Data 中时才这样写。

## Reused public data

```text
This study used publicly available [dataset name/type] from [Repository or source], available under
[DOI/accession/stable URL]. We used [version/release/date accessed, if relevant]. No new primary
[data type] data were generated for this part of the analysis.
```

中文对应:使用公开数据库时,需要写清数据库名、版本/发布日期/访问日期和编号,并引用数据集。

## Mixed generated and reused data

```text
Data generated in this study are available in [Repository] under [DOI/accession]. Public datasets
reused in the analysis were obtained from [source 1, identifier/version] and [source 2,
identifier/version]. Source data for [figures/tables] are provided in [location].
```

中文对应:自己产生的数据和复用的公开数据要分开写,避免让读者误以为所有数据都是本研究产生。

## Controlled-access human or sensitive data

```text
The [data type] data supporting this study are not publicly available because [privacy, consent,
legal, ethical or security reason]. A metadata record is available at [repository/accession, if
available]. Qualified researchers may request access from [data access committee/institutional
office/repository procedure] at [contact or URL]. Access requires [ethics approval/data-use
agreement/other conditions] and will be reviewed according to [policy or committee name].
```

中文对应:涉及人类参与者、隐私或伦理限制时,不能只写“因隐私不可公开”;还要写申请路径和审核条件。

## Third-party or licensed data

```text
The [data type/name] data used in this study were obtained from [third-party provider] under
licence and are not publicly redistributable by the authors. Requests for access should be directed
to [provider/contact/URL]. Derived data that can be shared are available in [repository] under
[DOI/accession], subject to [licence or restriction].
```

中文对应:第三方授权数据不能由作者重新分发时,要说明数据所有者和读者应向谁申请。

## Commercially restricted data

```text
The [data type] data are subject to commercial restrictions and cannot be made publicly available.
Requests for access may be directed to [company/data owner/contact or URL] and are subject to
[approval/licence/payment/confidentiality terms]. The authors provide [summary statistics,
metadata, synthetic data, or source data] in [location] to support interpretation of the results.
```

中文对应:企业或商业数据不可公开时,需要说明商业限制、申请对象,以及是否有汇总数据或元数据可公开。

## Embargoed data

Use only when the repository supports embargo and the journal permits it.

```text
The [data type] data have been deposited in [Repository] under [DOI/accession] and are under
embargo until [date/event]. Reviewers can access the data using [private reviewer link or
repository access route]. The data will become publicly available at [DOI/accession] when the
embargo ends.
```

中文对应:如果数据暂时不公开,必须已有仓库记录、审稿访问方式和明确解封时间或条件。

## Request-based access with justified restriction

```text
The [data type] data are not publicly available because [specific reason]. Requests for access may
be sent to [institutional group/contact route], and will be considered for [eligible purpose/users]
subject to [approval, agreement, or legal condition]. [Public metadata/aggregate data/source data]
are available at [location].
```

中文对应:“合理请求”只有在说明原因、接收机构、审核条件和可公开元数据后才可接受。

## No datasets generated or analysed

Use sparingly.

```text
No datasets were generated or analysed during the current study.
```

中文对应:只有确实没有生成或分析任何数据时才能使用,经验研究通常不适用。

For theory papers, be more specific:

```text
This work is theoretical and does not generate or analyse empirical datasets.
```

## Anti-patterns to revise

| Weak wording | Why it fails | Stronger move |
|---|---|---|
| Data are available upon request. | No reason, route, eligibility, or durability. | Add restriction reason, responsible access body, conditions, and metadata. |
| Data are available from the corresponding author on reasonable request. | Often a literal translation of "可向通讯作者合理索取"; not durable or specific enough. | Use an institutional/repository access route and define review conditions. |
| Data will be uploaded after acceptance. | No current repository or durable identifier. | Deposit before submission or provide a private reviewer link. |
| All data are in the manuscript. | Often false for figures/statistics. | Name exact source data, supplementary files, and omitted raw data. |
| Data are proprietary. | Does not say who controls access. | Name owner/provider and access route. |
| N/A. | Nature-style instructions usually require an explanation. | State why no datasets were generated or analysed. |

## Audit questions

- Which result would fail if this dataset were unavailable?
- Is the route durable beyond the corresponding author's current email address?
- Can a reader tell what each identifier contains?
- Are restrictions specific enough for an editor to judge them?
- Are reused datasets cited, not merely mentioned?
SKILL.md
---
name: nature-data
description: >-
  Prepare, audit, or revise Nature-ready Data Availability statements, data repository plans,
  dataset citations, and FAIR metadata checklists for manuscripts. Use when the user asks about
  Nature data availability, research data sharing, repository selection, accession numbers,
  restricted or sensitive data, source data, supplementary datasets, DataCite-style dataset
  references, FAIR metadata for academic publication, or Chinese-to-English data availability
  wording for Chinese-speaking authors preparing Nature-family submissions.
  Also trigger on general academic-writing data needs even without the word "Nature", such as
  writing a data availability statement for any journal, code/data sharing sections, repository
  selection while writing a paper, and Chinese phrasings like 数据可用性声明、数据可用性、
  数据共享、代码可用性、学术写作数据声明、写数据声明、数据存放、数据仓库选择.
metadata:
  author: Yuan1z skill, refactored into static/dynamic layers
---

# Nature Data Availability — Router

This skill is split into two layers:

- A **static layer** under `static/` that holds versioned, reusable content fragments (the default stance and source hierarchy, the Chinese-user operating mode, and the workflow with output format).
- A **dynamic layer** (this file plus `manifest.yaml`) that loads the core every time and reaches for the deeper policy/repository/FAIR references only when a step needs them.

Do not try to apply the data-availability logic from memory or from this router. Always load fragments from disk as described below.

## Routing protocol

Follow these four steps every time the skill is invoked.

### 1. Load the manifest and the core layer

Read [manifest.yaml](manifest.yaml). Then read every file listed under `always_load`:

- `static/core/stance.md` — what the data-availability package is, the default stance, and the source hierarchy.
- `static/core/chinese-mode.md` — how to operate when the user writes in Chinese (accept Chinese, draft English, convert terms precisely).
- `static/core/workflow.md` — the eight-step workflow and the output format.

### 2. No content axis — confirm journal and language inline

Unlike nature-writing or nature-figure, nature-data has no fragment axis. Its variation is handled at runtime, not by loading different content bodies:

- **journal/article type** — if journal-specific instructions conflict with this skill, follow the journal.
- **access route** — each dataset is classified into one route (public repository, controlled access, within paper, reused public, third-party restricted, justified request, or not applicable).
- **user language** — if the user writes Chinese, follow `core/chinese-mode.md` and add the 中文核对 block.

### 3. Run the workflow

Follow the eight-step workflow in `core/workflow.md`: identify the journal, inventory every supporting dataset, classify each into one access route, choose repository and identifier strategy before drafting, draft the statement with explicit dataset-to-location mapping, add formal dataset citations, run the FAIR/metadata audit, and return ready-to-paste text plus unresolved fields.

Do not invent DOIs, accession numbers, repository names, licences, embargo dates, ethics approvals, access committees, or data-use conditions. Flag "available upon request" as weak unless there is a specific legal, ethical, commercial, or third-party restriction.

### 4. Reach for references only when needed

The files under `references/` are deep references, not defaults. Open them on demand per the `references.on_demand` table in the manifest — for example `references/policy-principles.md` for the governing rules and edge cases, `references/repository-and-identifiers.md` for repository/accession/DOI choices, `references/statement-patterns.md` for ready-to-adapt statements, `references/fair-metadata-checklist.md` for the FAIR audit, `references/chinese-author-alignment.md` for Chinese wording, and `references/source-basis.md` to justify a rule with its official source.

When the target is the flagship journal Nature, also open
`references/nature-article-requirements.md` for statement placement,
mandatory-deposition routing, central-code review access, materials and
structure-file checks.

When the target is Nature Machine Intelligence, open
`../nature-shared/journal-formats/nature-machine-intelligence.md`. Enforce a
Data Availability statement and a separate `Code availability` section after
it and before references; check reviewer access, precise restrictions,
repository/identifier quality and the Software Submission Checklist for newly
developed central code.

## Why this split

- The static layer is versioned and reviewable; the core stays small for a normal statement.
- The dynamic layer keeps each invocation cheap: the policy, repository, and FAIR depth load only when a step needs them.
- The router itself is short on purpose. Update fragments and references, not this file, when adding scope.
- This structure mirrors `nature-writing`, `nature-polishing`, `nature-reader`, `nature-paper2ppt`, `nature-figure`, `nature-citation`, and `nature-response`.
static/core/chinese-mode.md
# Chinese-user operating mode

When the user writes in Chinese, provides a Chinese manuscript note, or asks for "中文对应", "中英对照", "数据可用性声明", "数据获取声明", "原始数据", "数据存储库", or "受限数据":

- Accept Chinese input naturally, but draft the final submission-ready statement in English unless the user explicitly asks for Chinese only.
- Preserve a short Chinese explanation of unresolved decisions when it helps the author act.
- Translate intent, not wording. Chinese phrases such as "可向通讯作者索取" are usually too vague for Nature-style English unless the restriction and access process are specified.
- Convert Chinese repository/status descriptions into precise publication terms:
  `数据可用性声明` -> `Data Availability`; `原始数据` -> `raw data`; `处理后数据` -> `processed data`; `源数据` -> `source data`; `补充材料` -> `Supplementary Information`; `受限数据` -> `restricted data`; `合理请求` -> `reasonable request`, only with reason and review route.
- Use `references/chinese-author-alignment.md` for Chinese terminology, common CN-to-EN failure modes, and bilingual intake questions.
static/core/stance.md
# Default stance and source hierarchy

Use this skill to turn a manuscript's supporting data into a transparent, Nature-ready data availability package: statement text, repository plan, dataset citations, and missing-information flags.

The governing policy layer is Springer Nature / Nature Portfolio data policy. The implementation layer is FAIR data practice and DataCite-style citation metadata.

## Default stance

- Treat the Data Availability statement as a link between the paper's claims and the evidence needed to inspect, reproduce, or reuse them.
- Do not invent DOIs, accession numbers, repository names, licences, embargo dates, ethics approvals, access committees, or data-use conditions.
- Prefer public, discipline-specific repositories. Use generalist or institutional repositories only when no suitable community repository exists.
- Describe both newly generated data and reused third-party data.
- If data cannot be openly shared, state why, who controls access, how requests are evaluated, and what metadata or representative data can still be public.
- Separate data, code, materials, and protocols unless the journal asks for a combined availability section.
- Keep this skill focused on availability and metadata. Do not rewrite methods, analyze statistics, or polish the manuscript unless the user asks for those tasks separately.
- Flag "available upon request" as weak unless there is a specific legal, ethical, commercial, or third-party restriction.

## Source hierarchy

Use sources in this order:

1. Target journal instructions and submission system requirements.
2. Nature Portfolio / Springer Nature data, code, materials, and reporting policies.
3. Repository-specific requirements and domain community standards.
4. FAIR principles and DataCite metadata practice.

If a policy detail may have changed, verify the current journal page before giving final submission advice.
static/core/workflow.md
# Workflow and output format

## Workflow

1. Identify the target journal and article type. If journal-specific instructions conflict with this skill, follow the journal.
2. Inventory every dataset needed to support the main and supplementary results: generated raw data, processed data, figure source data, secondary data, software outputs, models, tables, images, and files underlying statistical analysis.
3. Classify each dataset into one access route: `public repository`, `controlled access repository`, `within paper or supplement`, `reused public source`, `third-party restricted`, `available on justified request`, or `not applicable`.
4. Choose repository and identifier strategy before drafting text. Prefer DOI, accession number, Handle, ARK, or stable repository record over personal websites and temporary cloud links.
5. Draft the Data Availability statement using explicit dataset-to-location mapping.
6. Add formal dataset citations for public data that support conclusions.
7. Run the FAIR and metadata audit before finalizing.
8. Return ready-to-paste statement text plus any unresolved fields the author must confirm.

## Output format

Unless the user asks for another format, return:

```text
Data Availability
[ready-to-paste statement]

Repository and citation actions
- [specific actions or "None"]

Missing information / risk flags
- [specific flags or "None"]

中文核对
- [用中文列出作者需要确认的字段或 "无"]
```

When auditing an existing statement, lead with blocking issues first, then provide a revised version.