yuan1z0825/nature-skills包含需要注意的行为
SKILL DETAIL
nature-paper-card
yuan1z0825/nature-skills/nature-paper-card
nature-paper-card 是一个用于构建源基深度阅读 Paper Card 的技能,适用于单篇科学论文、预印本、PDF、DOI、arXiv 页面、出版商文章或粘贴的论文文本。当用户请求 Paper Card、深度阅读文献卡片、单篇论文深度分析、模块化分析、实验到主张的证据链、结论边界审计、批判性分析、知识连接或候选研究思路时使用。该技能生成固定的 01-16 章节,涵盖文献定位、研究问题、背景路线、痛点、核心洞见、方法与模块逻辑、关键公式、实验到主张的证据、结论边界、作者声明的局限性、批判性分析、所学知识、知识连接和可测试的研究思路。不适用于全文双语翻译、正式同行评审报告、批量文献监测、学术英语收集、理解测验或公共文章写作。
安装量 · 145查看来源
Installation
npx skills add https://github.com/yuan1z0825/nature-skills --skill nature-paper-card
技能文件
SKILL.md
最近同步 · 2026年8月27日
agents/openai.yaml›
interface:
display_name: "Nature Paper Card"
short_description: "Build evidence-grounded deep-reading paper cards"
default_prompt: "Use $nature-paper-card to turn this paper into a source-grounded Sections 01-16 deep-reading Paper Card."
evals/evals.json›
{
"skill_name": "nature-paper-card",
"evals": [
{
"id": 1,
"prompt": "Turn this methods paper into a complete Paper Card. Explain why each module exists, which ablation supports it, where the experimental conclusions stop, and propose testable research ideas without claiming novelty.",
"expected_output": "A paper-card.md with Sections 01-16, source-grounded method and module analysis, experiment-to-claim mapping, bounded conclusions, separated author limitations and Agent analysis, and hypothesis-labelled research candidates."
},
{
"id": 2,
"prompt": "I currently have only the abstract and metadata. Generate a preliminary Paper Card without guessing any unavailable information.",
"expected_output": "A visibly Partial Paper Card that identifies Abstract-only coverage, preserves all Sections 01-16 headings, fills only supported fields, and marks unseen methods, figures, experiments, limitations, and ideas as not assessable or provisional."
},
{
"id": 3,
"prompt": "Generate a Paper Card for this experimental science paper. Separate author-acknowledged limitations from your critical analysis, and determine whether the evidence supports association, necessity, sufficiency, or mechanism.",
"expected_output": "A discovery-oriented Paper Card with causal-strength discipline, figure/evidence pointers, author-stated limitations isolated in Section 12, and falsifiable Agent analysis in Section 13."
},
{
"id": 4,
"prompt": "Write the Paper Card in Chinese for this algorithm paper, which also introduces a major training dataset. Distinguish PDF page indices from printed page labels and audit both the system and dataset contributions.",
"expected_output": "A Chinese Paper Card using a primary methods lens plus a secondary resource lens, localized headings, unambiguous PDF page pointers, complete figure/table inventory, and separate card-completeness and context-verification statuses."
},
{
"id": 5,
"prompt": "The bundled preparation script cannot extract reliable page indices from this paper, but section headings, Figure 2, Table 1, and the Methods text are readable. Generate the Paper Card without writing any replacement Python.",
"expected_output": "A structure-grounded Paper Card that declares the canonical locator mode, uses section/figure/table/equation pointers without page numbers, records the preparation failure, and runs the bundled auditor without a source bundle."
},
{
"id": 6,
"prompt": "Only the abstract and bibliographic metadata are available. Generate the most honest Paper Card possible without inventing methods, experiments, figures, tables, limitations, or page numbers.",
"expected_output": "A source-limited partial Paper Card using only Abstract or Metadata pointers, explicit not-assessable markers, no page citations, and a structure-only audit warning that source-inventory coverage could not be checked."
}
]
}
manifest.yaml›
name: nature-paper-card
version: 1.2.0
description: >
Declarative routing for a source-grounded Sections 01-16 Paper Card. The paper_type
axis changes the analytical checks while the final card structure remains
stable across disciplines.
always_load:
- ../nature-shared/core/terminology-ledger.md
- static/core/principles.md
- static/core/workflow.md
- static/core/output-contract.md
axes:
paper_type:
detect: |
Select one primary lens from the dominant argument. Select one secondary
lens only when a second contribution carries substantial independent
evidence and would otherwise be under-audited:
methods - algorithm, model, tool, framework, or system contribution
discovery - empirical discovery or biological/physical mechanism
resource - dataset, benchmark, atlas, database, or community resource
clinical - patient, population, diagnostic, prognostic, or intervention study
materials - material design, synthesis, characterization, or engineering performance
review - review, perspective, commentary, or meta-analysis
Classify by evidence logic rather than field label. Load no more than two
values. A methods paper with a substantial dataset contribution may use
methods as primary and resource as secondary.
values:
methods: static/fragments/paper_type/methods.md
discovery: static/fragments/paper_type/discovery.md
resource: static/fragments/paper_type/resource.md
clinical: static/fragments/paper_type/clinical.md
materials: static/fragments/paper_type/materials.md
review: static/fragments/paper_type/review.md
default: methods
multi: true
references:
on_demand:
- condition: assigning provenance labels and separating source facts, external facts, analysis, and hypotheses
path: references/evidence-and-provenance.md
- condition: writing or auditing the exact Sections 01-16 Markdown structure
path: references/card-schema.md
- condition: generating, evaluating, or qualifying candidate research ideas
path: references/research-idea-gates.md
quality_tools:
prepare_script: scripts/prepare_paper.py
prepare_use_when: processing a PDF or nature-reader source-map JSON
prepare_output: source_bundle.json
audit_script: scripts/audit_paper_card.py
audit_use_when: after the first complete Paper Card draft and after every corrective revision
audit_output: audit-report.json
script_resolution: resolve relative to the loaded SKILL.md directory, never the user working directory
replacement_scripts: forbidden during normal Paper Card generation
locator_modes:
- page-grounded
- structure-grounded
- source-limited
README_EN.md›
# `nature-paper-card` Skill
[中文说明](README.md)
`nature-paper-card` deeply reads one scientific paper and produces a source-grounded, reviewable Sections 01–16 Paper Card. It focuses on the research question, method logic, experiment-to-claim evidence, conclusion boundaries, critical analysis, and testable research ideas instead of translating the abstract.
## What To Use It For
- Deep-read one paper through a consistent structure rather than produce only a summary.
- Trace central conclusions to figures, tables, equations, experiments, and ablations.
- Separate author statements, external facts, Agent analysis, and research hypotheses.
- Examine conclusion boundaries, author-stated limitations, and unresolved questions.
- Derive falsifiable and actionable candidate research ideas from the evidence chain.
## Typical Requests
- "Use `nature-paper-card` to deep-read this PDF and generate a complete Paper Card."
- "Analyze the method modules, essential equations, and experiment-to-claim evidence chain."
- "What does this paper actually demonstrate, and what does it not demonstrate?"
- "Propose testable follow-up directions based on the paper's limitations."
## What You Need To Provide
- A paper PDF, DOI, arXiv page, publisher article, pasted text, or an existing `nature-reader` source map.
- When relevant, the output language, output directory, and questions that need special attention.
- The skill can work from an abstract or partial material, but it will explicitly mark sections that cannot be assessed.
## Workflow
1. Run the bundled `prepare_paper.py` to prepare source material instead of writing a temporary PDF extraction script.
2. Select `page-grounded`, `structure-grounded`, or `source-limited` mode according to evidence reliability.
3. Select a methods, discovery, resource, clinical, materials, or review lens from the paper's argument structure.
4. Build an evidence inventory and claim–evidence matrix before drafting the fixed Sections 01–16.
5. Run the bundled `audit_paper_card.py` to check structure, locators, and evidence grounding.
## Outputs
- `paper-card.md`: the fixed Sections 01–16 deep-reading Paper Card.
- `source_bundle.json`: normalized source material from a PDF or source map.
- `audit-report.json`: structure, locator, and evidence-grounding audit results.
- Optional `rendered-pages/`: rendered PDF pages for visual inspection.
## Runtime and Dependencies
- Python 3 is required.
- PDF processing uses the bundled scripts and PDF libraries available in the current environment.
- When PDF page indices are reliable, the card uses both PDF-page and structural locators. If page extraction fails, it falls back to structural locators without inventing page numbers.
- Missing or invalid source-map page locators remain explicitly unlocated with a warning; they are never assigned to PDF page 1.
- External search is limited to field-history checks, knowledge connections, bibliographic verification, or an explicitly requested novelty check.
## Tutorial
See the complete [English tutorial](../../docs/nature-paper-card-tutorial_EN.md) for copyable prompts, inputs, the three locator modes, output files, and acceptance checks.
Minimal invocation:
```text
Use nature-paper-card to deep-read this paper and generate an English Paper Card.
Focus on the method modules, decisive experiments, conclusion boundaries,
and testable follow-up ideas.
```
## Boundaries
- Process one paper at a time; do not perform batch literature monitoring.
- Do not produce full-text bilingual translation; use `nature-reader` for that output.
- Do not produce a formal peer-review report; use `nature-reviewer` for reviewer-style assessment.
- Do not turn the Paper Card into a public article or add Sections 17 and 18.
- When source material is insufficient, write `Not assessable` rather than invent unseen experiments or page numbers.
## Related Skills
- `nature-reader`: generate bilingual full-text Markdown, figure-text alignment, and a stable source map.
- `nature-academic-search`: verify field history, external knowledge connections, or related work.
- `nature-reviewer`: produce a formal reviewer-style assessment.
- `nature-literature-pipeline`: discover, screen, and deliver papers in batches.
- `nature-paper2ppt`: convert paper content into presentation slides.
README.md›
# `nature-paper-card` 技能
[English](README_EN.md)
`nature-paper-card` 用于精读一篇科研论文,并生成有来源约束、可复核的 01–16 节 Paper Card;它强调研究问题、方法逻辑、实验—结论证据链、结论边界、批判性分析和可检验研究想法,而不是把摘要翻译成中文。
## 适合用它做什么
- 对单篇论文进行结构化精读,而不是只生成摘要。
- 追踪核心结论对应的图、表、公式、实验和消融证据。
- 区分作者陈述、外部事实、Agent 分析和研究假设。
- 检查论文的结论边界、作者自述局限和未解决问题。
- 从证据链出发提出可证伪、可执行的候选研究想法。
## 典型请求
- “使用 `nature-paper-card` 精读这篇 PDF,生成完整 Paper Card。”
- “分析这篇论文的方法模块、关键公式和实验—结论证据链。”
- “这篇论文真正证明了什么、没有证明什么?”
- “基于论文中的限制,提出几个可以验证的后续研究方向。”
## 你需要提供
- 一篇论文的 PDF、DOI、arXiv 页面、出版社文章、粘贴文本,或已有的 `nature-reader` source map。
- 如有需要,说明输出语言、输出目录和希望重点审查的问题。
- 若只有摘要或局部材料,Skill 仍可运行,但会明确标记无法判断的部分。
## 工作方式
1. 使用内置 `prepare_paper.py` 准备来源材料,不临时重写 PDF 提取脚本。
2. 根据证据可靠性选择 `page-grounded`、`structure-grounded` 或 `source-limited` 模式。
3. 按论文的论证结构选择方法、发现、资源、临床、材料或综述分析镜头。
4. 先建立证据清单和 claim–evidence matrix,再生成固定的 01–16 节 Paper Card。
5. 使用内置 `audit_paper_card.py` 检查结构、来源定位和证据约束。
## 产出
- `paper-card.md`:固定 01–16 节的深度 Paper Card。
- `source_bundle.json`:PDF 或 source map 的标准化来源包。
- `audit-report.json`:结构、定位和证据约束审计结果。
- 可选的 `rendered-pages/`:需要视觉核对时渲染的 PDF 页面。
## 运行和依赖
- 需要 Python 3。
- PDF 处理优先使用 Skill 自带脚本和当前环境中可用的 PDF 库。
- PDF 页码可靠时同时使用 PDF 页码和结构定位;页码提取失败时自动降级为结构定位,不伪造页码。
- source map 中缺失或非法的页码会保留为未定位块并给出警告,不会自动归入 PDF 第 1 页。
- 外部检索只用于文献背景核验、知识连接、书目信息确认或用户明确要求的新颖性检查。
## 使用教程
完整的可复制教程见[中文教程](../../docs/nature-paper-card-tutorial.md),包括输入、触发方式、三种来源定位模式、输出文件和结果验收方法。
最小调用示例:
```text
使用 nature-paper-card 精读这篇论文,生成中文 Paper Card。
重点检查方法模块、关键实验、结论边界和可验证的后续研究想法。
```
## 边界
- 一次处理一篇论文,不负责批量文献监控。
- 不生成全文双语翻译;需要全文阅读材料时使用 `nature-reader`。
- 不生成正式同行评审报告;需要审稿人视角时使用 `nature-reviewer`。
- 不把 Paper Card 改写成公众号文章,也不生成第 17、18 节。
- 来源不足时明确写 `Not assessable`,不会补写不可见的实验或页码。
## 相关技能
- `nature-reader`:生成全文双语 Markdown、图文对应和稳定 source map。
- `nature-academic-search`:核验领域历史、外部知识连接或相关工作。
- `nature-reviewer`:生成正式的审稿人式评审报告。
- `nature-literature-pipeline`:批量发现、筛选和推送论文。
- `nature-paper2ppt`:把论文内容转换成汇报幻灯片。
references/card-schema.md›
# Paper Card schema
## Contents
- [01 Basic Information](#01-basic-information)
- [02 One-Sentence Summary](#02-one-sentence-summary)
- [03 Research Question](#03-research-question)
- [04 Research Background and Development Path](#04-research-background-and-development-path)
- [05 Core Pain Points Identified by the Paper](#05-core-pain-points-identified-by-the-paper)
- [06 Core Idea](#06-core-idea)
- [07 Method Overview](#07-method-overview)
- [08 Core Module Breakdown](#08-core-module-breakdown)
- [09 Essential Formulas and Symbols](#09-essential-formulas-and-symbols)
- [10 Experimental Design and Evidence Chain](#10-experimental-design-and-evidence-chain)
- [11 Correct Interpretation of the Conclusions](#11-correct-interpretation-of-the-conclusions)
- [12 Limitations Explicitly Acknowledged by the Authors](#12-limitations-explicitly-acknowledged-by-the-authors)
- [13 Critical Analysis](#13-critical-analysis)
- [14 Knowledge Learned](#14-knowledge-learned)
- [15 Connections to Existing Knowledge](#15-connections-to-existing-knowledge)
- [16 Research Ideas](#16-research-ideas)
Use all headings in this order. Translate the headings and tables into the user's language when needed. Keep a section concise when the paper does not require extensive treatment.
## 01 Basic Information
Record title, authors and affiliations, venue, year, paper type, field, keywords, DOI or arXiv identifier, code, dataset, reading date, and the paper's position in the user's research direction. Mark unavailable fields.
## 02 One-Sentence Summary
Answer in one sentence: what problem, what approach, through what mechanism, and what bounded result. Avoid promotional adjectives.
## 03 Research Question
State:
- the concrete problem;
- why it matters;
- why existing approaches are insufficient;
- one precise `Can ... ?` research question when appropriate.
## 04 Research Background and Development Path
Map stages, representative approaches, advantages, limitations, and the paper's claimed position. Label whether this path is externally verified or only framed by the paper.
## 05 Core Pain Points Identified by the Paper
Use:
| Pain point | Manifestation | Cause or author explanation | Evidence from the paper |
|---|---|---|---|
Do not turn an author explanation into an established root cause without evidence.
## 06 Core Idea
Separate:
1. surface method;
2. core insight;
3. possible general lesson.
Label the general lesson `[Analysis]`.
## 07 Method Overview
Record input, output, modules, training, tools, feedback loop, and assumptions. Add a text flow from input to output.
## 08 Core Module Breakdown
Use:
| Module | Function | Why it is needed | Input and output | Supporting evidence | Known or expected effect of removal |
|---|---|---|---|---|---|
Distinguish measured ablation effects from expected effects.
## 09 Essential Formulas and Symbols
Include only formulas essential to understanding. For each, give the formula, symbol meanings, purpose, intuition, and source pointer. Use `Not applicable` when appropriate.
## 10 Experimental Design and Evidence Chain
First record datasets or population, scale, metrics, baselines, budget, backbone or instrument, oracle inputs, and evaluation protocol.
Then use:
| Experiment | Claim tested | Comparison and conditions | Result | Supported conclusion | Unsupported stronger conclusion | Source |
|---|---|---|---|---|---|---|
## 11 Correct Interpretation of the Conclusions
Audit task scope, oracle or ground-truth inputs, end-to-end status, compute cost, historical-data dependence, model dependence, hardest cases, population or domain boundary, and uncertainty. End with a bounded restatement of the main result.
## 12 Limitations Explicitly Acknowledged by the Authors
Include only author-acknowledged limitations:
| Limitation | Specific manifestation | Future direction proposed by the authors | Source |
|---|---|---|---|
If none are explicit, write `No explicit author-acknowledged limitation was found in the supplied source.` Do not fill the table. If useful, add a separate `Related constraints noted by the authors` subsection, clearly stating that these are not presented by the paper as formal limitations.
## 13 Critical Analysis
Use:
| `[Analysis]` Observation | Potential issue or alternative explanation | Why it matters | How to test it | Basis |
|---|---|---|---|---|
Include only specific, falsifiable concerns. Do not imitate a formal reviewer report.
## 14 Knowledge Learned
Extract transferable concepts, methods, formulas, and experimental designs. If not supplied by the user, title the subsection `Agent-derived knowledge candidates`.
## 15 Connections to Existing Knowledge
Connect the paper to verified external literature, user-provided knowledge, or clearly marked candidate directions. Cover similarities, combinations, conflicts, and transferable domains only when supported.
## 16 Research Ideas
For each candidate include:
- name;
- originating limitation or observation;
- core hypothesis;
- delta from the paper;
- initial method;
- validation;
- possible failure modes;
- innovation status: `unverified`, `partially checked`, or `prior-art checked`.
Title Agent-generated content `Agent-derived research candidates`, not `My research ideas`.
references/evidence-and-provenance.md›
# Evidence and provenance rules
## Labels
Use the smallest applicable label:
| Label | Meaning | Required support |
|---|---|---|
| `[Paper]` | The paper explicitly reports or claims this | Page/section/figure/table/equation/block pointer |
| `[External]` | A source outside the paper supports this | Direct citation or URL |
| `[Analysis]` | The Agent infers this from identified evidence | Reasoning plus relevant source pointers |
| `[Hypothesis]` | A testable but unverified explanation or idea | Proposed test and possible falsifier |
| `[User]` | The user supplied this judgment or connection | User statement or linked user knowledge |
## Source-pointer format
Prefer:
```text
[Paper: PDF p. 3 (printed p. 24959), Fig. 2]
[Paper: PDF p. 6, Table 2]
[Paper: Eq. 6]
[Paper: S041-S044]
```
For PDFs in page-grounded mode, always write `PDF p. N` for the file page index. Store printed page labels in the source bundle and optionally declare their range once in the card header; do not require them in every pointer. Use only pointers available in the source. Never invent line numbers. For HTML without pages, use section and figure/table IDs. For abstract-only material, use `[Paper: Abstract]` and do not infer unseen methods or results.
## Locator modes
### Page-grounded
Use when the bundled preparation script validates PDF page indices:
```text
[Paper: PDF p. 3, Figure 2]
[Paper: PDF p. 7, Table 4]
```
Printed page labels may be stored in the source bundle or declared once in the card header. Do not repeat them in every pointer unless the user asks.
### Structure-grounded
Use when page indices are unreliable but document structure is reliable:
```text
[Paper: Methodology, Symbolic Parsing]
[Paper: Figure 2]
[Paper: Table 4]
[Paper: Equation 3]
```
Do not include any page number.
### Source-limited
Use when only limited source material is reliable:
```text
[Paper: Abstract]
[Paper: Metadata]
[Paper: User-provided excerpt]
```
Do not infer unseen methods, experiments, figures, tables, equations, or limitations.
## Claim strength
Use verbs that match evidence:
- `reports`, `observes`, `is associated with`;
- `supports`, `is consistent with`;
- `suggests`, `does not distinguish`;
- `demonstrates necessity` only with a valid intervention or ablation;
- `demonstrates sufficiency` only with evidence that establishes sufficiency;
- `causes` only when the design supports causal inference.
## Paper framing versus verification
Prefix an unverified field history with:
```text
[Paper-framed; external verification not performed]
```
Do not call a method first, novel, state of the art, or unprecedented solely because the authors do.
## Contradictions
When prose, tables, figures, or supplements disagree:
1. record both values and their locations;
2. do not silently choose one;
3. mark the affected conclusion uncertain;
4. state what would resolve the discrepancy.
references/research-idea-gates.md›
# Research idea gates
Apply all gates before including an idea in Section 16.
## 1. Traceability gate
Tie the idea to a specific limitation, failure case, assumption, contradiction, or unexplained result. Include its source pointer.
## 2. Hypothesis gate
State a proposition that could be false. Reject vague ideas such as "improve efficiency" or "use a larger model".
## 3. Delta gate
State exactly what changes relative to the paper:
- representation;
- mechanism;
- data;
- objective;
- feedback;
- evaluation;
- deployment constraint.
## 4. Validation gate
Specify:
- comparison;
- controlled variables;
- metric;
- dataset or population;
- expected observation;
- falsifying result;
- compute or experimental budget when relevant.
## 5. Failure gate
Give at least two plausible reasons the idea may fail. Include attribution error, confounding, instability, cost, data scarcity, or invalid assumptions when applicable.
## 6. Novelty-language gate
Use:
- `candidate idea`;
- `hypothesis`;
- `possible extension`;
- `prior-art search required`.
Do not use `novel`, `first`, `unexplored`, or a numeric innovation score unless a dedicated prior-art search has been performed and cited.
scripts/audit_paper_card.py›
#!/usr/bin/env python3
"""Audit a Paper Card against its source bundle and output contract."""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any
SECTION_RE = re.compile(r"^##\s+(\d{2})\b.*$", re.M)
PAPER_POINTER_RE = re.compile(r"\[Paper:\s*([^\]]+)\]")
FORBIDDEN_SECTION_RE = re.compile(r"^##\s+(17|18)\b", re.M)
OUT_OF_SCOPE_TERMS = [
"academic english",
"self-quiz",
"comprehension quiz",
"public article",
"social-media copy",
"\u5b66\u672f\u82f1\u6587",
"\u81ea\u6d4b",
"\u516c\u4f17\u53f7",
]
IDEA_STATUS_TERMS = [
"innovation status",
"\u521b\u65b0\u72b6\u6001",
]
IDEA_VALIDATION_TERMS = [
"validation",
"how to test",
"\u5982\u4f55\u9a8c\u8bc1",
]
IDEA_FAILURE_TERMS = [
"failure mode",
"possible failure",
"\u53ef\u80fd\u5931\u8d25",
]
LOCATOR_MODES = ("page-grounded", "structure-grounded", "source-limited")
SOURCE_LIMITED_LOCATORS = ("abstract", "metadata", "user-provided excerpt")
def finding(level: str, code: str, message: str, **details: Any) -> dict[str, Any]:
item: dict[str, Any] = {"level": level, "code": code, "message": message}
if details:
item["details"] = details
return item
def contains_any(text: str, terms: list[str]) -> bool:
lower = text.lower()
return any(term.lower() in lower for term in terms)
def section_text(card: str, number: str) -> str:
match = re.search(
rf"^##\s+{re.escape(number)}\b[^\n]*\n(.*?)(?=^##\s+\d{{2}}\b|\Z)",
card,
re.M | re.S,
)
return match.group(1) if match else ""
def evidence_item_mentioned(card: str, item_id: str) -> bool:
match = re.match(r"^(Figure|Table|Equation)\s+(\d+[A-Za-z]?)$", item_id)
if not match:
return item_id.lower() in card.lower()
kind, number = match.groups()
aliases = {
"Figure": [
f"Figure {number}",
f"Fig. {number}",
f"\u56fe{number}",
f"\u56fe {number}",
],
"Table": [
f"Table {number}",
f"\u8868{number}",
f"\u8868 {number}",
],
"Equation": [
f"Equation {number}",
f"Eq. {number}",
f"Eq.{number}",
f"\u516c\u5f0f{number}",
f"\u516c\u5f0f {number}",
],
}
lower = card.lower()
return any(alias.lower() in lower for alias in aliases[kind])
def audit(
card: str,
bundle: dict[str, Any] | None,
locator_mode: str,
) -> dict[str, Any]:
findings: list[dict[str, Any]] = []
sections = SECTION_RE.findall(card)
expected = [f"{number:02d}" for number in range(1, 17)]
if sections == expected:
findings.append(finding("pass", "sections", "Sections 01-16 are present in order."))
else:
findings.append(
finding(
"error",
"sections",
"Sections 01-16 are missing, duplicated, or out of order.",
expected=expected,
found=sections,
)
)
forbidden_sections = FORBIDDEN_SECTION_RE.findall(card)
if forbidden_sections:
findings.append(
finding("error", "forbidden_sections", "Sections 17 or 18 must not be present.")
)
else:
findings.append(finding("pass", "forbidden_sections", "No Sections 17 or 18 found."))
preamble = card.split("## 01", 1)[0]
metadata_lines = [line for line in preamble.splitlines() if line.lstrip().startswith(">")]
if len(metadata_lines) >= 7:
findings.append(
finding("pass", "status_header", "The evidence-status header has at least seven fields.")
)
else:
findings.append(
finding(
"error",
"status_header",
"The evidence-status header must contain at least seven blockquote fields.",
found=len(metadata_lines),
)
)
if locator_mode in preamble.lower():
findings.append(
finding(
"pass",
"locator_mode_declaration",
f"The card declares locator mode: {locator_mode}.",
)
)
else:
findings.append(
finding(
"error",
"locator_mode_declaration",
f"The card header must declare canonical locator mode: {locator_mode}.",
)
)
pointers = PAPER_POINTER_RE.findall(card)
if pointers:
findings.append(
finding("pass", "paper_pointers", f"Found {len(pointers)} paper source pointers.")
)
else:
findings.append(finding("error", "paper_pointers", "No [Paper: ...] pointers found."))
page_pointers = [
pointer
for pointer in pointers
if "PDF p." in pointer or "PDF pp." in pointer
]
ambiguous = [
pointer
for pointer in pointers
if re.search(r"\bpp?\.", pointer)
and "PDF p." not in pointer
and "PDF pp." not in pointer
]
if locator_mode == "page-grounded":
if bundle is None:
findings.append(
finding(
"error",
"bundle_required",
"Page-grounded mode requires a source bundle.",
)
)
if not page_pointers:
findings.append(
finding(
"error",
"page_grounding",
"Page-grounded mode requires at least one PDF page pointer.",
)
)
elif ambiguous:
findings.append(
finding(
"error",
"ambiguous_pages",
"Page pointers must distinguish PDF page indices.",
pointers=ambiguous,
)
)
else:
findings.append(
finding("pass", "page_grounding", "PDF page pointers are explicit and unambiguous.")
)
else:
if page_pointers or ambiguous:
findings.append(
finding(
"error",
"page_pointers_forbidden",
f"{locator_mode} mode must not emit page-number citations.",
pointers=page_pointers + ambiguous,
)
)
else:
findings.append(
finding(
"pass",
"page_pointers_forbidden",
f"{locator_mode} mode contains no page-number citations.",
)
)
inventory = bundle.get("evidence_inventory", {}) if bundle else {}
if bundle is not None:
for key, label in (("figures", "figure"), ("tables", "table"), ("equations", "equation")):
items = inventory.get(key, []) if isinstance(inventory, dict) else []
missing = [
item.get("id", "")
for item in items
if not evidence_item_mentioned(card, item.get("id", ""))
]
if missing:
findings.append(
finding(
"error",
f"{key}_coverage",
f"Not every main {label} from the source bundle appears in the card.",
missing=missing,
)
)
else:
findings.append(
finding(
"pass",
f"{key}_coverage",
f"All {len(items)} inventoried {key} appear in the card.",
)
)
else:
findings.append(
finding(
"warning",
"source_inventory_unavailable",
"No source bundle was supplied; figure, table, and equation coverage was not audited.",
)
)
if locator_mode == "source-limited" and pointers:
invalid_limited = [
pointer
for pointer in pointers
if not any(token in pointer.lower() for token in SOURCE_LIMITED_LOCATORS)
]
if invalid_limited:
findings.append(
finding(
"error",
"source_limited_scope",
"Source-limited pointers must use only Abstract, Metadata, or User-provided excerpt.",
pointers=invalid_limited,
)
)
else:
findings.append(
finding(
"pass",
"source_limited_scope",
"All source-limited pointers use allowed source scopes.",
)
)
section_12 = section_text(card, "12")
section_13 = section_text(card, "13")
if section_12 and section_13:
findings.append(
finding("pass", "limitations_split", "Sections 12 and 13 are both present.")
)
section_16 = section_text(card, "16")
if section_16:
checks = {
"idea_status": contains_any(section_16, IDEA_STATUS_TERMS),
"idea_validation": contains_any(section_16, IDEA_VALIDATION_TERMS),
"idea_failure": contains_any(section_16, IDEA_FAILURE_TERMS),
}
for code, passed in checks.items():
findings.append(
finding(
"pass" if passed else "warning",
code,
(
f"Section 16 includes the required {code.replace('idea_', '')} field."
if passed
else f"Section 16 may be missing the {code.replace('idea_', '')} field."
),
)
)
out_of_scope = [term for term in OUT_OF_SCOPE_TERMS if term.lower() in card.lower()]
if out_of_scope:
findings.append(
finding(
"error",
"out_of_scope",
"The card contains content excluded by the skill scope.",
terms=out_of_scope,
)
)
else:
findings.append(finding("pass", "out_of_scope", "No excluded content detected."))
errors = sum(item["level"] == "error" for item in findings)
warnings = sum(item["level"] == "warning" for item in findings)
passes = sum(item["level"] == "pass" for item in findings)
return {
"schema_version": "1.0",
"summary": {
"status": "fail" if errors else ("pass_with_warnings" if warnings else "pass"),
"passes": passes,
"warnings": warnings,
"errors": errors,
},
"metrics": {
"locator_mode": locator_mode,
"sections_found": sections,
"paper_pointer_count": len(pointers),
"figures_in_bundle": len(inventory.get("figures", [])),
"tables_in_bundle": len(inventory.get("tables", [])),
"equations_in_bundle": len(inventory.get("equations", [])),
},
"findings": findings,
}
def print_report(report: dict[str, Any]) -> None:
summary = report["summary"]
print(
f"Audit status: {summary['status']} "
f"(passes={summary['passes']}, warnings={summary['warnings']}, errors={summary['errors']})"
)
for item in report["findings"]:
print(f"{item['level'].upper():7} {item['code']}: {item['message']}")
details = item.get("details")
if details:
print(f" {json.dumps(details, ensure_ascii=False)}")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Audit a Paper Card and its locator mode.")
parser.add_argument("--card", type=Path, required=True, help="Paper Card Markdown file.")
parser.add_argument("--bundle", type=Path, help="Optional source_bundle.json.")
parser.add_argument(
"--locator-mode",
required=True,
choices=LOCATOR_MODES,
help="Grounding mode declared by the Paper Card.",
)
parser.add_argument("--report", type=Path, help="Optional JSON audit report.")
return parser.parse_args()
def main() -> int:
args = parse_args()
if not args.card.is_file():
print("ERROR: --card must point to an existing file.", file=sys.stderr)
return 2
if args.bundle and not args.bundle.is_file():
print("ERROR: --bundle does not exist.", file=sys.stderr)
return 2
if args.locator_mode == "page-grounded" and not args.bundle:
print("ERROR: page-grounded mode requires --bundle.", file=sys.stderr)
return 2
try:
card = args.card.read_text(encoding="utf-8")
bundle = json.loads(args.bundle.read_text(encoding="utf-8")) if args.bundle else None
except (OSError, UnicodeDecodeError, json.JSONDecodeError) as exc:
print(f"ERROR: {exc}", file=sys.stderr)
return 2
report = audit(card, bundle, args.locator_mode)
print_report(report)
if args.report:
output = args.report.resolve()
output.parent.mkdir(parents=True, exist_ok=True)
output.write_text(json.dumps(report, ensure_ascii=False, indent=2), encoding="utf-8")
print(f"Wrote audit report: {output}")
return 1 if report["summary"]["errors"] else 0
if __name__ == "__main__":
raise SystemExit(main())
scripts/prepare_paper.py›
#!/usr/bin/env python3
"""Prepare a stable source bundle for the nature-paper-card workflow."""
from __future__ import annotations
import argparse
import hashlib
import json
import re
import sys
from collections import defaultdict
from pathlib import Path
from typing import Any
SECTION_NAMES = {
"abstract",
"introduction",
"background",
"related work",
"method",
"methods",
"methodology",
"experiments",
"experimental setup",
"experimental results",
"results",
"method analysis",
"discussion",
"limitations",
"conclusion",
"conclusions",
"acknowledgments",
"acknowledgements",
"references",
"appendix",
}
CAPTION_RE = re.compile(r"^\s*(Figure|Fig\.|Table)\s+(\d+[A-Za-z]?)\s*:\s*(.*)$", re.I)
EQUATION_RE = re.compile(r"\((\d{1,2})\)\s*$")
PRINTED_PAGE_RE = re.compile(r"^\s*(\d{4,6})\s*$", re.M)
PAGE_LOCATOR_FIELDS = ("page", "page_number", "pdf_page")
POSITIVE_INTEGER_RE = re.compile(r"^[1-9]\d*$")
def sha256_file(path: Path) -> str:
digest = hashlib.sha256()
with path.open("rb") as handle:
for chunk in iter(lambda: handle.read(1024 * 1024), b""):
digest.update(chunk)
return digest.hexdigest()
def clean_text(text: str) -> str:
text = text.replace("\r\n", "\n").replace("\r", "\n")
return "\n".join(line.rstrip() for line in text.splitlines()).strip()
def printed_page_label(text: str) -> str | None:
candidates = PRINTED_PAGE_RE.findall(text)
return candidates[-1] if candidates else None
def heading_candidates(
text: str, pdf_page: int | None, printed_page: str | None
) -> list[dict[str, Any]]:
headings: list[dict[str, Any]] = []
for line in text.splitlines():
value = " ".join(line.split()).strip(" :")
lower = value.lower()
numbered = re.match(r"^(?:\d+(?:\.\d+)*)\s+([A-Z][A-Za-z0-9 ,/&-]{2,80})$", value)
if lower in SECTION_NAMES or numbered:
headings.append(
{
"title": value,
"pdf_page": pdf_page,
"printed_page": printed_page,
}
)
return headings
def collect_caption(lines: list[str], start: int) -> str:
match = CAPTION_RE.match(lines[start])
if not match:
return ""
parts = [match.group(3).strip()]
for line in lines[start + 1 : start + 4]:
value = " ".join(line.split())
if not value or CAPTION_RE.match(value) or value.lower() in SECTION_NAMES:
break
if re.match(r"^[A-Z][A-Za-z ]{2,40}$", value) and not value.endswith("."):
break
parts.append(value)
return " ".join(part for part in parts if part)
def evidence_from_pages(pages: list[dict[str, Any]]) -> dict[str, list[dict[str, Any]]]:
figures: list[dict[str, Any]] = []
tables: list[dict[str, Any]] = []
equations: list[dict[str, Any]] = []
seen: set[tuple[str, str]] = set()
for page in pages:
lines = page["text"].splitlines()
for index, line in enumerate(lines):
caption_match = CAPTION_RE.match(line)
if caption_match:
kind = caption_match.group(1).lower()
number = caption_match.group(2)
canonical = "Table" if kind == "table" else "Figure"
key = (canonical, number)
if key not in seen:
seen.add(key)
item = {
"id": f"{canonical} {number}",
"caption": collect_caption(lines, index),
"pdf_page": page["pdf_page"],
"printed_page": page.get("printed_page"),
}
(tables if canonical == "Table" else figures).append(item)
equation_match = EQUATION_RE.search(line)
if equation_match:
number = equation_match.group(1)
key = ("Equation", number)
if key not in seen:
seen.add(key)
context_start = max(0, index - 3)
equations.append(
{
"id": f"Equation {number}",
"context": " ".join(lines[context_start : index + 1]).strip(),
"pdf_page": page["pdf_page"],
"printed_page": page.get("printed_page"),
}
)
def natural_key(item: dict[str, Any]) -> tuple[int, str]:
match = re.match(r"\D+(\d+)(.*)", item["id"])
return (int(match.group(1)), match.group(2)) if match else (999999, item["id"])
return {
"figures": sorted(figures, key=natural_key),
"tables": sorted(tables, key=natural_key),
"equations": sorted(equations, key=natural_key),
}
def prepare_pdf(path: Path, render_dir: Path | None, render_scale: float) -> dict[str, Any]:
try:
import fitz
except ImportError as exc:
raise RuntimeError(
"PyMuPDF is required for PDF input. Install it with: python -m pip install pymupdf"
) from exc
document = fitz.open(path)
pages: list[dict[str, Any]] = []
sections: list[dict[str, Any]] = []
if render_dir:
render_dir.mkdir(parents=True, exist_ok=True)
for index, page in enumerate(document):
text = clean_text(page.get_text("text"))
pdf_page = index + 1
printed_page = printed_page_label(text)
page_record = {
"pdf_page": pdf_page,
"printed_page": printed_page,
"text": text,
"character_count": len(text),
}
pages.append(page_record)
sections.extend(heading_candidates(text, pdf_page, printed_page))
if render_dir:
matrix = fitz.Matrix(render_scale, render_scale)
pixmap = page.get_pixmap(matrix=matrix, alpha=False)
pixmap.save(render_dir / f"page-{pdf_page:03d}.png")
metadata = {key: value for key, value in document.metadata.items() if value not in (None, "")}
evidence = evidence_from_pages(pages)
return {
"schema_version": "1.0",
"source_type": "pdf",
"source_path": str(path.resolve()),
"source_sha256": sha256_file(path),
"metadata": metadata,
"page_count": document.page_count,
"pages": pages,
"sections": sections,
"evidence_inventory": evidence,
"rendered_pages_dir": str(render_dir.resolve()) if render_dir else None,
"extraction": {
"engine": "PyMuPDF",
"visual_pages_rendered": bool(render_dir),
"confidence": "high" if all(page["character_count"] > 50 for page in pages) else "mixed",
},
}
def flatten_source_map(data: Any) -> list[dict[str, Any]]:
if isinstance(data, list):
return [item for item in data if isinstance(item, dict)]
if not isinstance(data, dict):
return []
for key in ("blocks", "source_blocks", "items", "entries"):
value = data.get(key)
if isinstance(value, list):
return [item for item in value if isinstance(item, dict)]
return []
def page_locator(block: dict[str, Any]) -> tuple[int | None, str, str | None, Any]:
"""Return a verified 1-based page or an explicit missing/invalid state."""
locator_field: str | None = None
locator_value: Any = None
for field in PAGE_LOCATOR_FIELDS:
if field not in block:
continue
value = block[field]
if value is None or (isinstance(value, str) and not value.strip()):
if locator_field is None:
locator_field = field
locator_value = value
continue
locator_field = field
locator_value = value
break
else:
return None, "missing", locator_field, locator_value
if isinstance(locator_value, bool):
return None, "invalid", locator_field, locator_value
if isinstance(locator_value, int):
return (
(locator_value, "verified", locator_field, locator_value)
if locator_value > 0
else (None, "invalid", locator_field, locator_value)
)
if isinstance(locator_value, float):
return (
(int(locator_value), "verified", locator_field, locator_value)
if locator_value.is_integer() and locator_value > 0
else (None, "invalid", locator_field, locator_value)
)
if isinstance(locator_value, str) and POSITIVE_INTEGER_RE.fullmatch(locator_value.strip()):
return int(locator_value.strip()), "verified", locator_field, locator_value
return None, "invalid", locator_field, locator_value
def source_block_text(block: dict[str, Any]) -> str:
text = (
block.get("original")
or block.get("source_text")
or block.get("text")
or block.get("content")
or ""
)
return clean_text(str(text)) if text else ""
def unlocated_block_record(
block: dict[str, Any], index: int, text: str, status: str, field: str | None, value: Any
) -> dict[str, Any]:
source_id = block.get("id") or block.get("block_id") or f"U{index:03d}"
return {
"id": str(source_id),
"block_type": block.get("type"),
"section": block.get("section") or block.get("section_title"),
"text": text,
"character_count": len(text),
"locator_status": status,
"locator_field": field,
"locator_value": value,
}
def prepare_source_map(path: Path) -> dict[str, Any]:
data = json.loads(path.read_text(encoding="utf-8"))
grouped: dict[int, list[str]] = defaultdict(list)
unlocated_blocks: list[dict[str, Any]] = []
verified_block_count = 0
missing_locator_count = 0
invalid_locator_count = 0
for index, block in enumerate(flatten_source_map(data), start=1):
text = source_block_text(block)
if not text:
continue
page_number, locator_status, locator_field, locator_value = page_locator(block)
if locator_status == "verified" and page_number is not None:
verified_block_count += 1
grouped[page_number].append(text)
continue
if locator_status == "missing":
missing_locator_count += 1
else:
invalid_locator_count += 1
unlocated_blocks.append(
unlocated_block_record(
block, index, text, locator_status, locator_field, locator_value
)
)
pages = [
{
"pdf_page": page_number,
"printed_page": None,
"text": clean_text("\n".join(grouped[page_number])),
"character_count": len(clean_text("\n".join(grouped[page_number]))),
}
for page_number in sorted(grouped)
]
structure_units = pages + [
{
"pdf_page": None,
"printed_page": None,
"text": block["text"],
"character_count": block["character_count"],
}
for block in unlocated_blocks
]
sections = [
heading
for page in structure_units
for heading in heading_candidates(page["text"], page["pdf_page"], page["printed_page"])
]
if unlocated_blocks and verified_block_count:
locator_reliability = "mixed"
elif unlocated_blocks:
locator_reliability = "unavailable"
else:
locator_reliability = "reliable"
return {
"schema_version": "1.1",
"source_type": "source_map",
"source_path": str(path.resolve()),
"source_sha256": sha256_file(path),
"metadata": data.get("metadata", {}) if isinstance(data, dict) else {},
"page_count": len(pages),
"pages": pages,
"unlocated_blocks": unlocated_blocks,
"locator_summary": {
"verified_page_blocks": verified_block_count,
"missing_page_locator_blocks": missing_locator_count,
"invalid_page_locator_blocks": invalid_locator_count,
"page_locator_reliability": locator_reliability,
},
"sections": sections,
"evidence_inventory": evidence_from_pages(structure_units),
"rendered_pages_dir": None,
"extraction": {
"engine": "source-map-normalizer",
"visual_pages_rendered": False,
"confidence": "mixed",
"page_locator_reliability": locator_reliability,
},
}
def validate_bundle(bundle: dict[str, Any]) -> dict[str, Any]:
errors: list[str] = []
warnings: list[str] = []
page_count = bundle.get("page_count", 0)
pages = bundle.get("pages", [])
metadata = bundle.get("metadata", {})
inventory = bundle.get("evidence_inventory", {})
source_type = bundle.get("source_type")
unlocated_blocks = bundle.get("unlocated_blocks", [])
locator_summary = bundle.get("locator_summary", {})
if not isinstance(page_count, int) or page_count < 0:
errors.append("page_count must be a non-negative integer")
if not isinstance(pages, list) or len(pages) != page_count:
errors.append("pages must contain exactly page_count records")
else:
located_text_available = any(page.get("character_count", 0) >= 50 for page in pages)
unlocated_text_available = isinstance(unlocated_blocks, list) and any(
block.get("character_count", 0) >= 50
for block in unlocated_blocks
if isinstance(block, dict)
)
if not located_text_available and not unlocated_text_available:
errors.append("no source block contains enough extractable text")
if source_type == "pdf" and page_count <= 0:
errors.append("PDF page_count must be a positive integer")
if source_type == "source_map":
if not isinstance(unlocated_blocks, list):
errors.append("unlocated_blocks must be an array")
if not isinstance(locator_summary, dict):
errors.append("locator_summary must be an object")
else:
missing_count = locator_summary.get("missing_page_locator_blocks", 0)
invalid_count = locator_summary.get("invalid_page_locator_blocks", 0)
if isinstance(missing_count, int) and missing_count > 0:
warnings.append(
f"{missing_count} source block(s) have no page locator; retained in unlocated_blocks"
)
if isinstance(invalid_count, int) and invalid_count > 0:
warnings.append(
f"{invalid_count} source block(s) have invalid page locators; retained in unlocated_blocks"
)
if not isinstance(metadata, dict) or not metadata.get("title"):
warnings.append("document title is unavailable")
if isinstance(inventory, dict):
evidence_count = sum(
len(inventory.get(key, [])) for key in ("figures", "tables", "equations")
)
if evidence_count == 0:
warnings.append("no figures, tables, or equations were detected")
else:
errors.append("evidence_inventory must be an object")
if errors:
locator_mode = "source-limited"
elif source_type == "pdf":
locator_mode = "page-grounded"
else:
locator_mode = "structure-grounded"
return {
"status": "invalid" if errors else ("valid_with_warnings" if warnings else "valid"),
"errors": errors,
"warnings": warnings,
"recommended_locator_mode": locator_mode,
"printed_page_labels_required": False,
}
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Prepare a normalized source bundle from a paper PDF or nature-reader JSON."
)
parser.add_argument("input", type=Path, help="Input PDF or source-map JSON.")
parser.add_argument("--output", type=Path, required=True, help="Output source_bundle.json.")
parser.add_argument("--render-dir", type=Path, help="Optional directory for rendered PDF pages.")
parser.add_argument("--render-scale", type=float, default=1.25, help="PDF render scale.")
return parser.parse_args()
def main() -> int:
args = parse_args()
source = args.input.resolve()
if not source.is_file():
print(f"ERROR: input file does not exist: {source}", file=sys.stderr)
return 2
try:
if source.suffix.lower() == ".pdf":
bundle = prepare_pdf(source, args.render_dir, args.render_scale)
elif source.suffix.lower() == ".json":
bundle = prepare_source_map(source)
else:
print("ERROR: input must be a PDF or JSON source map.", file=sys.stderr)
return 2
except (RuntimeError, ValueError, json.JSONDecodeError) as exc:
print(f"ERROR: {exc}", file=sys.stderr)
return 2
bundle["validation"] = validate_bundle(bundle)
output = args.output.resolve()
output.parent.mkdir(parents=True, exist_ok=True)
output.write_text(json.dumps(bundle, ensure_ascii=False, indent=2), encoding="utf-8")
inventory = bundle["evidence_inventory"]
print(f"Wrote source bundle: {output}")
print(f"Pages: {bundle['page_count']}")
print(f"Sections: {len(bundle['sections'])}")
print(f"Figures: {len(inventory['figures'])}")
print(f"Tables: {len(inventory['tables'])}")
print(f"Equations: {len(inventory['equations'])}")
validation = bundle["validation"]
print(f"Bundle validation: {validation['status']}")
print(f"Recommended locator mode: {validation['recommended_locator_mode']}")
for warning in validation["warnings"]:
print(f"WARNING: {warning}")
for error in validation["errors"]:
print(f"ERROR: {error}", file=sys.stderr)
return 1 if validation["errors"] else 0
if __name__ == "__main__":
raise SystemExit(main())
SKILL.md›
---
name: nature-paper-card
description: Build a source-grounded deep-reading Paper Card for one scientific paper, preprint, PDF, DOI, arXiv page, publisher article, or pasted paper text. Use when the user asks for a Paper Card, deep-reading literature card, single-paper deep analysis, module-by-module analysis, experiment-to-claim evidence chain, conclusion-boundary audit, critical analysis, knowledge connections, or candidate research ideas. Produce the fixed Sections 01-16 covering bibliographic position, research question, background route, pain point, core insight, method and module logic, essential formulas, experiment-to-claim evidence, conclusion boundaries, author-stated limitations, critical analysis, learned knowledge, knowledge connections, and testable research ideas. Do not use for full-paper bilingual translation, formal peer-review reports, batch literature monitoring, academic-English collection, comprehension quizzes, or public-article writing.
---
# Nature Paper Card - Router
Use this skill to turn one paper into an evidence-grounded research card, not a translated abstract, generic summary, reviewer report, or publication article.
The skill uses:
- a static core under `static/core/` for principles, workflow, and the fixed output contract;
- one paper-type fragment under `static/fragments/paper_type/`;
- on-demand references for evidence labels, the exact card schema, and research-idea checks.
## Routing protocol
Follow these steps every time.
### 1. Load the manifest and core layer
Read [manifest.yaml](manifest.yaml), then read every file under `always_load`. Do not generate the card from this router alone.
### 2. Establish the source boundary
Identify which material is available:
- full paper with figures and tables;
- paper text without reliable layout;
- abstract or metadata only;
- an existing `nature-reader` artifact with stable source IDs.
Prefer an existing `nature-reader` artifact when supplied. Do not repeat full bilingual translation or figure extraction. If only partial material is available, create a visibly partial card and mark every unsupported section `Not assessable from supplied material`.
For a PDF or `nature-reader` source-map JSON, the bundled script is mandatory.
1. Resolve `SKILL_DIR` as the directory containing this loaded `SKILL.md`.
2. Verify `SKILL_DIR/scripts/prepare_paper.py` exists.
3. Run exactly the bundled script by its resolved path:
```text
python "SKILL_DIR/scripts/prepare_paper.py" INPUT \
--output WORKDIR/source_bundle.json
```
Add `--render-dir WORKDIR/rendered-pages` when visual page review is needed. Inspect the script exit code and the bundle validation block before drafting.
For source-map input, also inspect `locator_summary` and `unlocated_blocks`. Only records under `pages` have verified positive PDF page locators. Missing or invalid page locators remain in `unlocated_blocks` with an explicit status and must be cited structurally, never as page 1.
Never write inline Python, a temporary extraction script, or a replacement script during a Paper Card run. Never patch the bundled scripts during a normal Paper Card run. Modify these scripts only when the user explicitly asks to develop, debug, or improve the skill itself.
Use this fixed locator state machine:
- `page-grounded`: the bundled script succeeds and validates reliable PDF page indices. Use PDF page plus structural locators. Printed page labels are optional metadata.
- `structure-grounded`: page extraction is unreliable, but reliable sections, figures, tables, equations, source blocks, or full text remain available. Do not emit page-number citations.
- `source-limited`: only an abstract, metadata, or user-provided excerpt is reliable. Do not emit page-number citations or infer unseen evidence.
If preparation fails, record the failure. Prefer an existing `nature-reader` source map or the environment PDF/OCR capability, but do not create a replacement script. Then enter the strongest supported fallback mode.
### 3. Classify the paper type
Use the manifest to choose one primary `paper_type` and, only for a genuinely hybrid paper, one secondary contribution lens:
- `methods`
- `discovery`
- `resource`
- `clinical`
- `materials`
- `review`
Load the primary fragment and no more than one secondary fragment. Classify by the paper's argument and evidence structure, not merely its discipline. State both selections before analysis. For example, an algorithm paper that also introduces a substantial dataset may use `methods` as the primary lens and `resource` as the secondary lens.
### 4. Build the evidence base before drafting
Build an internal evidence inventory before drafting. At minimum, enumerate:
- bibliographic metadata and access status;
- research question and claimed contribution;
- method components, assumptions, and data flow;
- every main figure, table, and essential equation with its argumentative role;
- experiments, baselines, metrics, ablations, and reported results;
- author-stated limitations;
- stable source pointers to pages, sections, equations, figures, tables, or `nature-reader` block IDs.
Then build a compact claim-evidence matrix linking each central claim to the evidence that supports it and to any unresolved gap.
Use external search only for Section 04, Section 15, bibliographic verification, or an explicit novelty check. Never present the paper's own related-work narrative as independently verified field history. Record whether the context mode is `paper-only`, `targeted external check`, or `externally verified`.
### 5. Generate the fixed Sections 01-16 Paper Card
Apply, in order:
1. core principles;
2. the selected paper-type fragment;
3. core workflow;
4. output contract.
Read [references/evidence-and-provenance.md](references/evidence-and-provenance.md) before making analytical or externally verified claims. Read [references/card-schema.md](references/card-schema.md) when drafting the final Markdown. Read [references/research-idea-gates.md](references/research-idea-gates.md) before writing Section 16.
Write a real Markdown artifact, defaulting to `paper-card.md`. Keep all 16 numbered sections in order, but write `Not applicable` or `Not assessable` instead of inventing content.
Match the user's language by default. The skill source and schema remain English, but localize the Paper Card headings and prose when the user writes in another language. Preserve canonical technical terms and formulas.
### 6. Run groundedness QA
Before delivery, resolve the bundled auditor from `SKILL_DIR`. In `page-grounded` mode, run:
```text
python "SKILL_DIR/scripts/audit_paper_card.py" \
--card WORKDIR/paper-card.md \
--bundle WORKDIR/source_bundle.json \
--locator-mode page-grounded \
--report WORKDIR/audit-report.json
```
In either fallback mode, run the same auditor without a bundle:
```text
python "SKILL_DIR/scripts/audit_paper_card.py" \
--card WORKDIR/paper-card.md \
--locator-mode structure-grounded-or-source-limited \
--report WORKDIR/audit-report.json
```
Replace the last value with the actual canonical mode. Treat audit errors as blockers. Review warnings with scientific judgment rather than suppressing them mechanically.
Also verify:
- numerical results match the source;
- the evidence inventory covers every main figure and table;
- every major method, result, boundary, and limitation has a source pointer;
- PDF page pointers distinguish PDF page index from printed page labels;
- author statements are separated from Agent analysis;
- external field-history claims have external citations or are marked unverified;
- proposed ideas are hypotheses, not novelty claims;
- Sections 17 and 18 do not exist;
- no academic-English collection, comprehension quiz, or public-article draft was added.
If the auditor itself cannot run, state that failure and manually apply only its documented checks. Do not write a substitute auditor.
## Script red lines
- Do not resolve bundled scripts relative to the user's current working directory.
- Do not write or execute inline Python as a substitute for either bundled script.
- Do not create `extract_pdf.py`, `parse_paper.py`, or another one-off replacement.
- Do not patch skill code during a normal Paper Card generation request.
- Do not fabricate page numbers when preparation fails.
- Do not remove all grounding in fallback mode; use structural locators or explicit source-scope locators.
## Relationship to adjacent skills
- Use `nature-reader` for full-text bilingual reading artifacts, extraction, and stable source maps.
- Use `nature-academic-search` when external literature is needed to verify field history or knowledge connections.
- Use `nature-reviewer` for formal reviewer-style manuscript assessment.
- Use `nature-literature-pipeline` for batch discovery and lightweight monitoring notes.
- Use `nature-paper2ppt` when the requested end product is a presentation.
Do not silently switch the requested Paper Card into any of these outputs.
static/core/output-contract.md›
# Output contract
Default deliverable:
```text
paper-card.md
```
Required working artifacts in page-grounded mode:
```text
source_bundle.json
audit-report.json
```
In fallback modes, `source_bundle.json` may be unavailable, but `audit-report.json` remains required when the bundled auditor can run. These are machine-readable support artifacts, not additional Paper Card sections.
For normalized source-map input, `pages` contains only blocks with verified positive page locators. Blocks with missing or invalid locators remain in `unlocated_blocks`; `locator_summary` reports the counts and reliability. A `null` page in structural evidence means unknown location and must never be rendered as page 1.
The file must contain:
- title and evidence-status header;
- input scope and paper-type classification;
- Sections 01-16 in order;
- source pointers for substantive paper-derived claims;
- provenance labels for external facts, analysis, hypotheses, and user-supplied judgments;
- explicit `Not applicable` or `Not assessable` markers where needed.
Do not create Sections 17 or 18.
## Evidence-status header
Include:
```markdown
> Source coverage: Full paper / Partial paper / Abstract only / Metadata only
> Extraction confidence: High / Mixed / Low
> Locator mode: page-grounded / structure-grounded / source-limited
> Primary analytical lens: ...
> Secondary analytical lens: None / ...
> Context verification: Paper-only / Targeted external check / Externally verified
> Card completeness: Complete relative to supplied source / Partial
```
`Complete relative to supplied source` does not imply that field history, novelty, or knowledge connections were independently verified.
Locator rules:
- `page-grounded`: use `[Paper: PDF p. N, Figure/Table/Equation/Section]`.
- `structure-grounded`: use `[Paper: Figure N]`, `[Paper: Table N]`, `[Paper: Equation N]`, or `[Paper: Section title]`; never include page numbers.
- `source-limited`: use `[Paper: Abstract]`, `[Paper: Metadata]`, or `[Paper: User-provided excerpt]`; mark unseen sections not assessable.
## Quality gates
Block final delivery when:
- a major reported number cannot be traced to the source;
- a main figure or table has not been inventoried;
- the card claims full coverage from partial input;
- author-stated limitations and Agent criticism are mixed;
- a research idea is described as novel without prior-art verification;
- irrelevant template content was invented to fill a section.
If a blocker cannot be resolved, downgrade the card to `Partial` and identify the exact gap.
Run the bundled auditor before final delivery. In fallback modes, run it without `--bundle` and expect an inventory-unavailable warning. A nonzero exit caused by audit errors blocks delivery. Audit success does not prove scientific correctness; manually review claims, evidence strength, and analytical conclusions.
static/core/principles.md›
# Core principles
Build a Paper Card that helps a researcher reconstruct the paper's reasoning and inspect its evidential limits.
## Non-negotiable stance
- Treat the paper as a set of claims supported by specific evidence, not as prose to summarize.
- Preserve the difference between what the authors report, what the evidence directly supports, and what the Agent infers.
- Prefer exact numbers, conditions, and source pointers over evaluative adjectives.
- Explain why a method component exists and what evidence supports its contribution.
- State what a result does not establish.
- Leave unsupported fields incomplete rather than filling the template speculatively.
- Keep recurring terminology consistent using the shared Terminology Ledger.
## Scope
Produce only one Paper Card with Sections 01-16. Do not add:
- full-paper bilingual translation;
- formal reviewer reports;
- academic-English phrase collection;
- comprehension tests or self-quiz;
- public-article or social-media copy.
## Language
Match the language of the user's request unless the user specifies another output language. Localize section headings, table headings, and explanatory prose. Keep method names, dataset names, symbols, formulas, and established technical terms in their canonical forms.
## Required provenance classes
Separate content into:
- `[Paper]` - directly reported in the paper;
- `[External]` - verified from sources outside the paper;
- `[Analysis]` - reasoned interpretation from the available evidence;
- `[Hypothesis]` - proposed explanation or research direction requiring testing;
- `[User]` - a judgment or connection supplied by the user.
Do not write Agent-generated analysis in the user's voice. Use `Agent-derived candidate` where a section asks for the user's knowledge, connection, or idea and the user has not supplied one.
## Applicability
Keep the 16 headings stable, but adapt their contents to the paper type. A review may not have system modules or ablations; a theoretical paper may not have datasets; a clinical paper may not have a conventional component-removal study. Use `Not applicable` with a reason instead of forcing a methods-paper template onto every paper.
For hybrid papers, apply one primary analytical lens and at most one secondary lens. The primary lens governs the card's main argument; the secondary lens adds missing checks without duplicating the whole card.
static/core/workflow.md›
# Paper Card workflow
## 1. Inspect and bound the source
Record the available source, completeness, figure/table availability, supplementary-material availability, extraction confidence, and page-number systems. Resolve DOI or arXiv metadata when possible. Never imply full-paper coverage from an abstract-only input.
For PDF or source-map JSON input, resolve and run the bundled `scripts/prepare_paper.py` before manual analysis. Do not create a replacement script. Inspect the generated metadata, page count, section list, evidence inventory, extraction confidence, and validation block. For source maps, inspect `locator_summary` and `unlocated_blocks`: missing or invalid locators remain structurally usable but have no PDF page. If rendered pages were requested, visually check the pages that contain central figures, tables, equations, or extraction anomalies.
Choose and record exactly one locator mode:
- `page-grounded` when PDF page indices are reliable;
- `structure-grounded` when only section, figure, table, equation, or source-block locators are reliable;
- `source-limited` when only abstract, metadata, or supplied excerpts are reliable.
Printed page labels never determine whether page-grounded mode is available.
## 2. Build an evidence inventory
Before interpreting the paper, inventory:
- every main section;
- every main figure and its argumentative role;
- every main table and its argumentative role;
- every essential equation;
- datasets, populations, or materials;
- baselines, metrics, evaluation settings, and uncertainty;
- explicit author limitation statements;
- supplements that are present or unavailable.
Do not draft the card until the inventory explains where the central method and result evidence lives.
## 3. Build a claim-evidence matrix
Create stable internal records for:
- central claim;
- source pointer;
- evidence type;
- exact result or quotation-free paraphrase;
- supported claim strength;
- unsupported stronger interpretation;
- confidence;
- contradictions or missing support.
In page-grounded mode, use PDF page index plus section, equation, figure, table, appendix, or existing source-block IDs. Printed page labels are optional. In structure-grounded mode, omit all page numbers and use the structural identifiers. In source-limited mode, use only `Abstract`, `Metadata`, or `User-provided excerpt`. Do not invent page or line numbers.
## 4. Reconstruct the paper's argument
Write the chain:
```text
problem
-> limitation of prior approaches
-> core insight
-> design choice
-> evidence required
-> experiment or analysis supplied
-> conclusion actually supported
```
Identify broken, weak, or untested links without turning the output into a reviewer report.
## 5. Analyze method and evidence
For each central component, answer:
- What does it do?
- Why is it needed?
- What are its inputs and outputs?
- What assumption does it introduce?
- Which experiment isolates its contribution?
- What remains unknown if it is removed or changed?
For each central experiment, answer:
- What claim is being tested?
- What is compared under which conditions?
- What metric and uncertainty are reported?
- What result was observed?
- What conclusion is justified?
- What stronger conclusion is not justified?
## 6. Add external context selectively
Use external sources for field history or knowledge connections only when accessible and useful. Cite them. If no external verification is performed, label the development route and novelty position as `paper-framed` or `unverified`.
Record one context mode:
- `paper-only` - no claims beyond the supplied paper are independently verified;
- `targeted external check` - selected metadata, history, or connections are verified;
- `externally verified` - Sections 04 and 15 receive a deliberate multi-source check.
Do not make context verification a prerequisite for a card that the user explicitly wants to remain source-only.
## 7. Separate limitations
Keep two sections distinct:
- Section 12 contains only limitations explicitly acknowledged by the authors.
- Section 13 contains Agent analysis, alternative explanations, missing controls, failure cases, and proposed validation tests.
Do not move Agent criticism into Section 12.
If the paper has no explicit limitations section or statement, say so. Do not promote incidental dataset descriptions, evaluation constraints, or implementation details into author-acknowledged limitations. Related author-noted constraints may be listed under a separately labelled subsection.
## 8. Generate knowledge and ideas
For Sections 14-16:
- extract transferable concepts, methods, formulas, and experimental designs;
- connect them to verified literature, user-provided knowledge, or clearly marked candidate directions;
- derive research ideas from a concrete limitation or unresolved observation;
- specify hypothesis, delta, validation, and failure modes;
- avoid claiming novelty without a dedicated search.
## 9. Write and audit the artifact
Write `paper-card.md` using the exact schema and the user's language. Check internal consistency across terminology, datasets, sample sizes, baselines, metrics, and numbers. Add a short evidence-status summary at the top.
Run the bundled `scripts/audit_paper_card.py` with the selected locator mode. Supply `source_bundle.json` in page-grounded mode. In fallback modes, the bundle is optional and unavailable source-inventory checks become warnings. Correct all errors, assess warnings, and rerun until no audit error remains. The script checks structure and traceability; it does not replace scientific judgment.
static/fragments/paper_type/clinical.md›
# Clinical / population paper
Emphasize the design-to-inference chain.
- Record population, setting, enrollment, exposure or intervention, comparator, outcomes, follow-up, missing data, and analysis set.
- Distinguish prespecified from exploratory outcomes and relative from absolute effects.
- Check randomization, blinding, confounding, multiplicity, calibration, external validation, attrition, adverse events, and clinical relevance.
- Separate statistical significance from effect size and practical benefit.
- Do not generalize beyond the studied population or convert observational association into causal effect.
- In Section 16, propose ethically feasible validation and state the target population explicitly.
static/fragments/paper_type/discovery.md›
# Discovery / mechanism paper
Emphasize the question-to-evidence chain.
- Identify the observation, causal or mechanistic claim, intervention, controls, replication, and alternative explanations.
- Separate association, necessity, sufficiency, and mechanism.
- Map each figure to the narrowing of the proposed mechanism.
- Check biological or physical scale, model-system limits, temporal order, measurement validity, and independent replication.
- Do not infer causality from correlation or a mechanism from a phenotype alone.
- In Section 16, propose discriminating experiments between competing explanations.
static/fragments/paper_type/materials.md›
# Materials / chemistry / engineering paper
Emphasize the design-to-performance and property-to-mechanism chains.
- Record composition, synthesis or fabrication, processing conditions, characterization, operating environment, and comparator.
- Link structure, property, mechanism, and performance without skipping evidential steps.
- Check batch variation, measurement uncertainty, durability, scale-up, realistic conditions, and normalization of performance metrics.
- Distinguish demonstrated mechanism from post-hoc interpretation.
- Compare against fair state-of-the-art baselines under matched conditions.
- In Section 16, propose mechanism-discriminating or scale-up experiments with failure criteria.
static/fragments/paper_type/methods.md›
# Methods / algorithm / system paper
Emphasize the problem-to-solution chain.
- Reconstruct inputs, outputs, modules, training requirements, external tools, feedback loops, and inference-time cost.
- Distinguish the core insight from the implementation bundle.
- Map each ablation to one component and check whether the comparison changes only that component.
- Check baseline fairness, backbone parity, data leakage, compute budget, oracle inputs, and end-to-end status.
- Treat a component's intended role as an author claim until an isolation experiment supports it.
- In Section 16, propose ideas that change a specific assumption, mechanism, or bottleneck and define a budget-fair validation.
- When a resource or dataset is a substantial secondary contribution, also load the `resource` lens instead of treating data volume as an unexplained implementation detail.
static/fragments/paper_type/resource.md›
# Resource / dataset / benchmark paper
Emphasize the construction-to-validation chain.
Use this fragment as a secondary lens when a methods paper introduces a dataset that materially drives the reported gains.
- Describe collection, inclusion/exclusion, annotation, preprocessing, splits, licensing, access, and intended use.
- Analyze representativeness, missingness, label quality, contamination, leakage, demographic or domain coverage, and update policy.
- Separate evidence that the resource is large from evidence that it is useful, valid, and generalizable.
- Inspect benchmark incentives and whether the metric measures the claimed capability.
- Treat downstream model improvements as conditional on evaluation design.
- In Section 16, propose validation, extension, stress-test, or governance ideas with measurable criteria.
static/fragments/paper_type/review.md›
# Review / perspective / meta-analysis
Emphasize the evidence-map chain.
- Identify scope, search or selection method, inclusion/exclusion rules, taxonomy, synthesis method, and intended contribution.
- For systematic reviews or meta-analyses, inspect protocol, heterogeneity, risk of bias, publication bias, and sensitivity analysis.
- For narrative reviews or perspectives, separate evidence synthesis from author framing and advocacy.
- Replace method-module ablation questions with taxonomy robustness, coverage, counterexamples, and omitted literatures.
- Do not treat citation volume as evidence quality.
- In Section 16, derive questions from unresolved disagreements, missing comparisons, or under-covered evidence rather than inventing a system module.
tests/test_prepare_paper.py›
#!/usr/bin/env python3
"""Regression tests for source-map page-locator normalization."""
from __future__ import annotations
import importlib.util
import json
import subprocess
import sys
import tempfile
import unittest
from pathlib import Path
sys.dont_write_bytecode = True
SCRIPT_PATH = Path(__file__).parents[1] / "scripts" / "prepare_paper.py"
SPEC = importlib.util.spec_from_file_location("nature_paper_card_prepare", SCRIPT_PATH)
assert SPEC and SPEC.loader
PREPARE = importlib.util.module_from_spec(SPEC)
SPEC.loader.exec_module(PREPARE)
LONG_TEXT = (
"This source block is deliberately longer than fifty characters so bundle "
"validation accepts it."
)
class SourceMapLocatorTests(unittest.TestCase):
def prepare(self, blocks: list[dict]) -> dict:
with tempfile.TemporaryDirectory() as directory:
source = Path(directory) / "source-map.json"
source.write_text(
json.dumps({"metadata": {"title": "Test paper"}, "blocks": blocks}),
encoding="utf-8",
)
bundle = PREPARE.prepare_source_map(source)
bundle["validation"] = PREPARE.validate_bundle(bundle)
return bundle
def test_verified_page_one_remains_page_one(self) -> None:
bundle = self.prepare([{"id": "S001", "page": 1, "text": LONG_TEXT}])
self.assertEqual([page["pdf_page"] for page in bundle["pages"]], [1])
self.assertEqual(bundle["unlocated_blocks"], [])
self.assertEqual(bundle["locator_summary"]["verified_page_blocks"], 1)
def test_missing_locator_does_not_become_page_one(self) -> None:
bundle = self.prepare([{"id": "S001", "text": LONG_TEXT}])
self.assertEqual(bundle["pages"], [])
self.assertEqual(bundle["page_count"], 0)
self.assertEqual(bundle["unlocated_blocks"][0]["locator_status"], "missing")
self.assertEqual(bundle["validation"]["recommended_locator_mode"], "structure-grounded")
self.assertEqual(bundle["validation"]["status"], "valid_with_warnings")
self.assertTrue(
any("no page locator" in warning for warning in bundle["validation"]["warnings"])
)
def test_non_numeric_locator_is_preserved_as_invalid(self) -> None:
bundle = self.prepare(
[{"id": "S001", "page": "not-a-page", "text": LONG_TEXT}]
)
block = bundle["unlocated_blocks"][0]
self.assertEqual(bundle["pages"], [])
self.assertEqual(block["locator_status"], "invalid")
self.assertEqual(block["locator_field"], "page")
self.assertEqual(block["locator_value"], "not-a-page")
def test_zero_locator_is_preserved_as_invalid(self) -> None:
bundle = self.prepare([{"id": "S001", "page": 0, "text": LONG_TEXT}])
block = bundle["unlocated_blocks"][0]
self.assertEqual(bundle["pages"], [])
self.assertEqual(block["locator_status"], "invalid")
self.assertEqual(block["locator_value"], 0)
def test_unknown_location_text_remains_available_for_structure_analysis(self) -> None:
equation_text = (
"Methods\nThe model minimizes the following objective while retaining "
"all structural evidence from the supplied source map. (2)"
)
bundle = self.prepare([{"id": "S009", "text": equation_text}])
self.assertEqual(bundle["unlocated_blocks"][0]["text"], equation_text)
self.assertEqual(bundle["sections"][0]["title"], "Methods")
self.assertIsNone(bundle["sections"][0]["pdf_page"])
self.assertEqual(bundle["evidence_inventory"]["equations"][0]["id"], "Equation 2")
self.assertIsNone(bundle["evidence_inventory"]["equations"][0]["pdf_page"])
def test_mixed_locators_never_merge_unknown_text_into_page_one(self) -> None:
bundle = self.prepare(
[
{"id": "S001", "page": 1, "text": LONG_TEXT},
{"id": "S002", "text": "Unlocated " + LONG_TEXT},
{"id": "S003", "page": "bad", "text": "Invalid " + LONG_TEXT},
]
)
self.assertEqual(len(bundle["pages"]), 1)
self.assertEqual(bundle["pages"][0]["text"], LONG_TEXT)
self.assertEqual(len(bundle["unlocated_blocks"]), 2)
self.assertEqual(
bundle["locator_summary"],
{
"verified_page_blocks": 1,
"missing_page_locator_blocks": 1,
"invalid_page_locator_blocks": 1,
"page_locator_reliability": "mixed",
},
)
def test_cli_preserves_an_invalid_locator_as_unknown(self) -> None:
with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
source = root / "source-map.json"
output = root / "source-bundle.json"
source.write_text(
json.dumps(
{
"metadata": {"title": "CLI test"},
"blocks": [{"page": "not-a-page", "text": LONG_TEXT}],
}
),
encoding="utf-8",
)
result = subprocess.run(
[sys.executable, str(SCRIPT_PATH), str(source), "--output", str(output)],
check=False,
capture_output=True,
text=True,
)
bundle = json.loads(output.read_text(encoding="utf-8"))
self.assertEqual(result.returncode, 0, result.stderr)
self.assertEqual(bundle["pages"], [])
self.assertEqual(bundle["unlocated_blocks"][0]["locator_status"], "invalid")
self.assertIn("invalid page locators", "\n".join(bundle["validation"]["warnings"]))
if __name__ == "__main__":
unittest.main()