返回 Skills 目錄
mirage-hq/tesseract已通過檢查

SKILL DETAIL

tesseract-motion

mirage-hq/tesseract/tesseract-motion

Create editable motion graphics locally in Tesseract, including full-frame animated scenes, typography, diagrams, lower thirds, and overlays on supplied footage. Use for focused motion-design work; use tesseract-video for assembling or revising a complete footage edit.

安裝量 · 597查看來源

Installation

npx skills add https://github.com/mirage-hq/tesseract --skill tesseract-motion

技能檔案

SKILL.md

最近同步 · 2026年9月23日

agents/openai.yaml
interface:
  display_name: "Tesseract: Motion Graphics"
  short_description: "Create editable motion graphics with Tesseract"
  icon_small: "./assets/icon.svg"
  icon_large: "./assets/icon.svg"
  default_prompt: "Use $tesseract-motion to create and preview editable motion graphics."
assets/icon.svg
<svg width="512" height="512" viewBox="0 0 512 512" fill="none" xmlns="http://www.w3.org/2000/svg">
<rect width="100%" height="100%" fill="#2A2D2C"/>
<g transform="translate(56 99.6618) scale(18.18181818)">
<path d="M0.00306569 8.59859C0.00170896 8.58085 0.000777255 8.56299 0.000380143 8.54499C-0.0477036 6.51024 4.47836 5.75717 6.40287 5.63163C6.43625 5.62944 6.43753 5.57365 6.40442 5.56884C3.98042 5.21632 2.36937 4.59234 2.36937 3.88126C2.36944 3.01774 4.74528 2.28259 8.06776 2.00618C8.10084 2.00343 8.10172 1.95205 8.06878 1.94808C6.51557 1.76153 5.48309 1.43103 5.48309 1.05447C5.48318 0.472103 7.95272 3.11982e-06 10.999 0H11.001C14.0473 1.94074e-06 16.5168 0.472103 16.5169 1.05447C16.5169 1.43103 15.4844 1.76153 13.9312 1.94808C13.8983 1.95206 13.8992 2.00342 13.9322 2.00618C17.2547 2.28259 19.6306 3.01774 19.6306 3.88126C19.6306 4.59234 18.0196 5.21632 15.5956 5.56884C15.5625 5.57366 15.5638 5.62943 15.5971 5.63163C17.5216 5.75716 22.0477 6.51024 21.9996 8.54499C21.9992 8.56299 21.9983 8.58085 21.9969 8.59859C21.9983 8.61633 21.9992 8.63419 21.9996 8.65219C22.0477 10.6869 17.5216 11.44 15.5971 11.5656C15.5638 11.5677 15.5625 11.6235 15.5956 11.6283C18.0196 11.9809 19.6306 12.6048 19.6306 13.3159C19.6306 14.1794 17.2547 14.9146 13.9322 15.191C13.8992 15.1938 13.8983 15.2451 13.9312 15.2491C15.4844 15.4357 16.5169 15.7661 16.5169 16.1427C16.5168 16.7251 14.0473 17.1972 11.001 17.1972H10.999C7.95272 17.1972 5.48318 16.7251 5.48309 16.1427C5.48309 15.7661 6.51557 15.4357 8.06878 15.2491C8.10172 15.2451 8.10084 15.1938 8.06776 15.191C4.74528 14.9146 2.36944 14.1794 2.36937 13.3159C2.36937 12.6048 3.98042 11.9809 6.40442 11.6283C6.43753 11.6235 6.43625 11.5677 6.40287 11.5656C4.47836 11.44 -0.0477036 10.6869 0.000380143 8.65219C0.000777255 8.63419 0.00170896 8.61633 0.00306569 8.59859Z" fill="url(#paint0_linear_532_1088)"/>
<defs>
<linearGradient id="paint0_linear_532_1088" x1="11" y1="0" x2="11" y2="17.8765" gradientUnits="userSpaceOnUse">
<stop stop-color="#5BF6BB"/>
<stop offset="0.173102" stop-color="#00DCFF"/>
<stop offset="0.620228" stop-color="#EFFBFF"/>
<stop offset="0.793276" stop-color="#FFCFCF"/>
<stop offset="1" stop-color="#F99F3F"/>
</linearGradient>
</defs>
</g>
</svg>
references/ad-hooks.md
# Build a visual hook for an ad

For a new paid-social, YouTube, or other video ad, deliberately design its opening few seconds. Derive the hook from the actual product, footage, audience, promise, and evidence. Make it visually compelling from the first frame, with a clear reason for the intended viewer to keep watching. Sound and transitions should strengthen that idea. A hook may be a powerful shot, an overlay, or a complete motion scene; it need not be an extra intro card.

Apply this when creating an ad or improving its opening. A narrow revision elsewhere does not require rebuilding the hook. Preserve an opening the user has explicitly chosen unless changing it is part of the request.

## Choose the idea before the treatment

Inspect the source material and brief. Identify the viewer's relevant problem/desire, the strongest truthful benefit or surprising moment, and where the body of the ad delivers on it. Use the most compelling source moment even if it occurs late in the footage, provided reordering does not distort what happened or what the speaker meant.

For a substantial new ad, consider a few distinct concepts—usually two or three is enough—then select and build the strongest within the authorized edit. Compare concepts by immediate clarity, relevance, visual interest, available evidence, fit with the brand/music, and the transition into the body. Different fonts or whooshes on the same opening are treatment variations, not different hook ideas. No separate approval is needed for each ordinary creative choice; render multiple variants when requested or when a close choice makes a short comparison useful.

| Hook approach | What the viewer sees | Use it when |
|---|---|---|
| Result first | The best real outcome or most impressive working moment, then its explanation. | The supplied result is immediately legible and credible. |
| Recognizable friction | A specific frustrating task, awkward moment, or bottleneck in action. | The audience can recognize its relevance without a long setup. |
| Visible contrast | An A/B, a matched comparison, or a transformation that makes the difference obvious. | Both sides are supported by the material; labels and timing make the comparison fair. |
| Unexpected perspective | A tight detail, unusual crop, scale shift, or reveal that resolves into the product/use case. | The surprise helps the viewer understand the offering instead of becoming unrelated spectacle. |
| A visual question | An incomplete action, interrupted process, or concealed detail whose answer begins to emerge quickly. | The ad actually supplies the answer; suspense does not postpone all meaning until later. |
| A distinctive human moment | A revealing expression, direct statement, or action already present in the footage, framed around its meaning. | The person and moment carry more interest than an effects sequence. |
| An idea made visible | Editable type, objects, paths, or supplied assets collide, sort, multiply, connect, or transform to explain the benefit. | Motion graphics communicate something the available footage cannot. |
| Evidence up close | An actual product interaction, useful detail, or supported demonstration becomes the opening focal point. | The evidence itself is compelling enough to carry attention. |

These are prompts, not an exhaustive library or a mandated style. Combine approaches only when the opening still has one clear idea. Avoid generic “stop scrolling” copy, unsupported superlatives, fabricated metrics/testimonials, unrelated shock footage, and mystery that never connects to the product. A stronger hook can come from better shot selection or a better opening line without adding more effects.

## Give the opening a small narrative

Choose its duration from the placement and material; roughly the first 2–5 seconds is a useful working range, not a universal rule. Start with meaningful visual information rather than black, a slow fade, or a logo-only wait. In the first few seconds establish the relevant tension/result, reveal enough to orient the viewer, and hand off naturally into the demonstration or explanation. Integrate brand/product recognition into the action when possible.

Keep one dominant focal point. Use concise text only where it adds context, with enough stable reading time at the actual display size. A caption, hook headline, UI screen, and animated label should not all compete. Fit the supplied brand and ad style: a direct creator-led opening and an elaborate launch-film opening require different treatments. A quiet but surprising image can work; maximum speed and volume are not the objective.

Match the placement:

- **Scrolling feeds, Reels, and Shorts:** make the first frame meaningful and the proposition visually understandable without audio; sound-on playback should add impact. Preserve UI safe areas and readable phone-scale text. TikTok recommends introducing the proposition within the first three seconds and developing the hook in the first six; treat these as planning guidance, not a requirement to spend six seconds on an intro. [TikTok creative guidance](https://ads.tiktok.com/resources/help/article/creative-best-practices?lang=en)
- **YouTube skippable in-stream:** give a clear reason to keep watching and establish product/brand relevance before the skip opportunity. Google documents skipping after five seconds; this is not the timing model for every YouTube placement. [Google ad formats](https://support.google.com/google-ads/answer/2375464?hl=en)
- **Very short ads:** let the hook and main benefit form the same concise action; do not spend most of the runtime on setup. Across formats, Google's creative guidance emphasizes entering the story quickly, early branding, and audio/text that reinforce rather than compete with the message. [Google ABCDs](https://support.google.com/google-ads/answer/14783551?hl=en)

Use the specified placement/delivery brief. If unspecified, make a sensible assumption from the request and state it in the edit plan; do not block creative work with a platform questionnaire.

## Time the hook to the actual music

When music is supplied or already generated, follow [waveform editing](waveform-editing.md) **before authoring the hook timing**. Inspect and audition the beginning of the music segment actually used by the edit. Zoom into its useful opening accents, pickup, first downbeat, build, dropout, or change in texture. The candidate detector helps locate events; confirm them by listening.

Record a compact cue map: `music source time → project time → visual event → SFX event`. Specify the actual reveal/arrival/impact, not just “sync to beat.” Include the selected music in-point and its project start so source, project, and layer clocks stay distinct.

For each chosen cue:

1. Set the perceptual visual event on that cue: a cut, fully exposed result, shutter reopening, key word settling, or meaningful motion arrival. Animate the lead-in before it and leave a useful hold afterward.
2. Compute SFX placement from its own perceptual event. A riser may start earlier so it resolves on the cue; a downlifter may begin on the cue and trail afterward; a shutter or ping aligns its audible transient. File starts and largest sample peaks are not automatically the right anchors.
3. Quantize picture events to the actual project frame rate, keeping the error within the nearest representable frame and aligning the audible event as closely as the runtime permits. Check the rendered result; matching JSON timestamps alone does not prove perceptual synchronization.

For example, if listening confirms an opening accent at **0.8 s in project time**, show a meaningful close-up from frame zero, begin the reveal shortly beforehand, expose the result at 0.8 s, and let the next shot explain it. This timing is illustrative; derive real values from the track. If the music starts softly or the first big hit comes late, create immediate visual interest and use an earlier subtle cue or motivated SFX. Do not make viewers wait through an empty intro for the drop. Preserve explicitly chosen music/in-points; do not silently replace or retime them. Without music, derive rhythm from speech, action, and selected sound accents; do not invent a soundtrack requirement.

## Use diverse sound and transition treatments

Choose by the hook's idea and physical motion. Possible treatments include a match cut, object wipe, shutter close/reopen, brief light sweep or leak, freeze-and-release, mask reveal, whip movement, split-screen change, directional blur, or a hard cut into an unexpected scale. Keep useful subject/product detail visible and resolve into readable content. A light leak or shutter should earn its place in this particular ad, not become the plugin's default opening.

Sounds can include bass downlifters, impacts with audible upper harmonics, pings, pops, clicks, camera shutters, paper/physical textures, or a short pause in the bed. Choose compatible cues and leave space for the music and voice. A subtle transient may beat a pile of risers, hits, and whooshes; a strong existing musical hit may need no additional accent. Check that the hook still communicates when muted and that low-frequency cues translate beyond headphones.

Use supplied/local audio or prepare and import procedural accents through [sound design](sound-design.md). The procedural set includes `downlifter`, `soft-impact`, `clear-ping`, `soft-pop`, and `dry-click`; it does not include a literal `camera-shutter` preset. Use an available local shutter asset or an intentional stylized accent rather than inventing a helper command. Follow [motion design](motion-design.md) and the actual schema to build transitions with editable shapes, masks, transforms, and supported effects. There is no assumed one-click `lightLeak` or `cameraShutter` field. Use native constructs or permitted local assets, with the timing kept editable.

## Review the opening as its own deliverable

Render the opening plus its transition into the body early, before polishing the entire ad. Inspect a filmstrip containing the first frame, setup, key reveal, readable hold, and the next scene. Around synchronized events, inspect neighboring frames and the rendered audio waveform; play the result at speed with sound. If timing is off, correct the responsible layer/cue and re-render the affected interval.

Then review it muted and at phone scale. Ask: Is there a clear focal point immediately? Why would this audience care? What benefit or question does the opening establish? Does the body deliver on it? Is the product/brand connection clear enough? Are text and motion readable? Do music and SFX strengthen the same event without masking words? Reject openings that are impressive but confusing, misleading, disconnected, or slow to reveal relevance.

Record the selected concept, source evidence, actual cue times, and what was inspected in the edit notes. Describe the intended mechanism rather than claiming guaranteed retention, conversion, or virality. If campaign results are later supplied, assess early retention alongside the campaign's downstream objective; attention alone does not establish a successful ad. No publishing, ad spend, or campaign changes are part of this hook-authoring workflow.
references/audio-and-timing.md
# Audio and timing

Read [waveform editing](waveform-editing.md) before timing an edit to music or trimming dialogue. Use its local `waveform` command for timestamped overview/zoom PNGs, inspect those views, and verify music events and word boundaries by listening. Recheck changed speech joins on the rendered edit clock before final mixing.

## Listen to the source before mixing

Probe for audio streams, then audition the dialogue, important production sounds, and music. Preserve speech intelligibility and meaningful foley. A clip with an audio stream is not necessarily usable sound. Add appropriate local sound design as described in [sound design](sound-design.md); do not invent narration or replace deliberately chosen music.

When a transcript is supplied, align cuts to its actual source timestamps. If there is no local transcription tool and speech cannot be understood reliably, ask for the transcript or flag the limitation. Never invent spoken words or caption timing.

For uneven, noisy, boxy, or reverberant dialogue, read [speech cleanup](speech-cleanup.md) before processing. It covers source selection, EQ, dynamics, de-essing, room-sound limits, and matched-loudness review of local derivatives.

## Select or make a local accent

```bash
python3 "$SKILL/scripts/tesseract_sound.py" list
python3 "$SKILL/scripts/tesseract_sound.py" make --preset air-whoosh --duration-ms 600 --seed 12 --output /absolute/path/edit/.tesseract-work/whoosh-v1.wav
```

Generate a fresh WAV with `make`; no preset WAV files or rendering binaries are bundled. The helper contains soft-tap, dry-click, soft-pop, clear-ping, air-whoosh, short-riser, downlifter, and soft-impact. Each is a small procedural accent, not a music track or recorded foley. Listen in context before choosing it; the default file peak is −12 dBFS and does not establish its perceived mix level.

The helper creates and inspects local files. Import a selected cue, music track,
or cleaned dialogue derivative with `project import-asset --kind audio` before
adding its editable Audio layer. See [media import](media-import.md) for commands.
Import packages the bytes; placement and mixing are separate checkout/commit edits.

Use the imported asset ID in an Audio layer, with `captionsEnabled: false` for music/effects. Keep the cue's source interval inside the actual file and its active interval inside the composition. A cue peaking 300 ms into the file should start 300 ms before the visual event it punctuates. Trim/fade a tail near the edit's end rather than unintentionally extending the project.

## Local mix

New Video layers need `volume: 1.0` to enable embedded source audio at unity gain; omitted/null volume disables it. Preserve existing gains on revisions. FX `Audio` layers can place music, a supplied sound effect, or dialogue independently. They have no visual transform. A minimal source carries `assetId`; source and active ranges are explicit.

```json
{
  "id": 90,
  "name": "Supplied music bed",
  "type": "Audio",
  "activeRange": {"start": 0, "duration": 8000},
  "sourceRange": {"start": 1500, "duration": 8000},
  "sourceIntrinsicDuration": 30000,
  "source": {"assetId": "local-music-id"},
  "volume": 0.2,
  "captionsEnabled": false
}
```

Use the schema to check additional source metadata or per-channel features. Do not infer all audio effects of a DAW from the presence of an Audio layer.

`volume` is linear: unity is `1.0`; `0.5` is approximately −6 dB. Convert dB with `10 ** (dB / 20)`. A fixed value such as `0.2` is a starting point, not a calibrated mix target. Loud source material needs different gain.

## Envelopes and ducking

For an FX Audio layer, target `propertyType: "volume"` using `layerTimeJsCode`. This is called `AudioVolume` in Rust, but the JSON wire name is `volume`, not `audioVolume`. Explicit ramps prevent clicks and allow a music bed to duck under dialogue. For example, fade in, hold quietly during speech, then lift:

```js
var t = input.time.seconds;
function lerp(a,b,p){return a+(b-a)*Math.max(0,Math.min(1,p));}
if(t < 0.2) return lerp(0,0.12,t/0.2);
if(t < 3.8) return 0.12;
if(t < 4.2) return lerp(0.12,0.30,(t-3.8)/0.4);
if(t < 7.5) return 0.30;
return lerp(0.30,0,(t-7.5)/0.5);
```

A `volume` animator supplies the **final linear layer gain**, overriding the static `volume`; it is not automatically multiplied by that static value. Include the chosen bed level in every returned value.

Use `setFxPropertyAnimator` for the script, or supported volume keyframe actions
from the installed action schema. Apply them with `tsrct project apply`. For a
bed at −14 dB that ducks by 6 dB, return the linear equivalent of −20 dB during
speech. Set times from the actual dialogue, not the illustrative script above.
If a root Audio layer starts at project time 5s, its local 950 ms duck occurs at
project time 5.95s; account for ancestor playback when nested.

For another duration, change the envelope accordingly. The times are owner-layer-local. Moving the Audio layer should move the envelope with it.

## Native controls versus local audio preparation

This package supports native Audio layers, source/active timing, linear gain, and time-varying volume. Use envelopes for predictable ducking, fades, and planned dropouts. This is **timed ducking**, not an audio-reactive sidechain compressor.

Do not invent native `compressor`, `sidechain`, `EQ`, `pan`, or `truePeakLimiter` fields. When the source needs EQ/compression, prepare a new **local audio derivative** with installed FFmpeg, keep the original source and processing recipe, and import it with `project import-asset --kind audio` and mix it through Tesseract. Do not flatten the editable project or quietly change the original media.

For example, a gentle music presence dip can be auditioned with:

```bash
ffmpeg -nostdin -n -i /absolute/path/music.wav -af "equalizer=f=2000:t=q:w=0.8:g=-3" -ar 48000 -c:a pcm_s24le /absolute/path/music-presence-v1.wav
```

Real signal-driven ducking, when needed, can be prepared with FFmpeg `sidechaincompress` using **aligned full-length** music and dialogue stems, both starting at the same edit zero. Its first input is the music to process; the second is the dialogue detector. Preserve the separate dialogue track when assembling the Tesseract mix. Adjust threshold/ratio from the source, check attack/release and actual reduction, and do not apply both the preprocessed duck and a second unplanned native duck. Misaligned, trimmed, or short detector stems need alignment/padding first. A source with music and voice already mixed together is not a clean dialogue detector.

Use `ffmpeg -h filter=<name>` to verify the locally installed filter/options before processing. Filter behavior and available controls are described in the [official FFmpeg filter reference](https://ffmpeg.org/ffmpeg-filters.html#sidechaincompress).

## Editing sound

- An audio prelap or tail uses independent placement and source selection. Do not detach dialogue without retaining lip sync where the speaker is visible.
- Avoid duplicate playback from both a Video layer and a copied Audio layer. If moving sound into independent layers, deliberately mute the corresponding original audio.
- Use whooshes/impacts only for a movement or transition that benefits from them. A sound on every text entrance quickly becomes distracting.
- Preserve the selected music track's identity. Do not replace, regenerate, or globally retime supplied music unless the user asks.
- Speed-changing a group also changes descendant audio clocks. Check source duration and pitch/time-stretch behavior in the actual render instead of assuming a silent visual retime.

## Review

Listen to the final muxed video, not just separate audio files. Check the head and tail, dialogue-to-music balance, abrupt gain edges, and source synchronization. Measure loudness/peaks locally with FFmpeg if needed; for example `ffmpeg -i edit.mp4 -af ebur128=peak=true -f null -`. Treat readings as measurements, not proof that speech is intelligible. Keep reasonable peak headroom and follow an explicit delivery spec when one is provided.

## Measure and check the final delivery

```bash
python3 "$SKILL/scripts/tesseract_sound.py" measure /absolute/path/edit-v1.mp4 --report /absolute/path/edit/.tesseract-work/checks/loudness-v1.json
python3 "$SKILL/scripts/tesseract_sound.py" mobile-preview /absolute/path/edit-v1.mp4 --output /absolute/path/edit/.tesseract-work/checks/mobile-v1.wav
```

`measure` reports the first audio stream's integrated LUFS, true peak in dBTP, and loudness range; it does not normalize or modify the input. The working range is −16 to −14 LUFS and true peak at most −1 dBTP, subject to the actual delivery spec. Short isolated effects or silence may not have meaningful integrated measurements. If the deliverable has several audio programs, measure the intended stream explicitly with FFmpeg.

For mastering, first adjust the native mix. If final normalization is needed, preserve the editable mix, use a measured two-pass `loudnorm` process on a derived deliverable, and record the recipe. Check the local [loudnorm options](https://ffmpeg.org/ffmpeg-filters.html#loudnorm): use the measured input loudness, range, peak, threshold, and target offset from pass one in pass two, then remeasure the encoded output. Do not normalize each individual cue to the full-program LUFS target or erase an intentional dropout to satisfy a meter.

`mobile-preview` creates a separate mono, 150 Hz–7 kHz listening copy. It is useful for noticing disappearing bass-dependent cues, phase cancellation, or poor voice balance, but it is not the final mix and does not replace listening on an actual phone.
references/capabilities.md
# Capability map for Tesseract

This skill uses the installed Tesseract CLI and its version-matched schemas.
Instructions describe intent; exact fields come from `project schema` (actions)
and `project schema --document` (editable JSON). A successful save is not proof
that every requested combination renders correctly; inspect the actual result.

| Area | Tesseract mechanism | Read |
|---|---|---|
| Footage editing | Imported Video layers; independent source/active ranges, source audio, transform and framing | [native authoring](native-authoring.md), [FX authoring](fx-authoring.md#local-footage) |
| Compositing | Native text, shapes, groups, media, masks, adjustments, and effects supported by the document schema | [FX authoring](fx-authoring.md) |
| Motion graphics | Editable keyframe actions, procedural scripts, and coordinated groups | [motion](motion.md), [motion design](motion-design.md) |
| Type | Local TTF/OTF/TTC import, editable text, paragraph layout and schema-supported text animators | [text](fx-authoring.md#text-with-an-available-font) |
| Captions | Editable phrase text layers with actual words/timing from supplied or verified material | [bottom captions](fx-authoring.md#bottom-captions) |
| Custom effects | Bounded premultiplied-alpha WGSL with declared parameters and layer texture inputs | [custom shader contract](custom-shader.md) |
| Audio | Embedded video sound; imported Audio layers, source timing, linear gain and envelopes | [audio and timing](audio-and-timing.md) |
| Local sound tools | Procedural WAV accents, source/edit waveform PNGs, loudness measurements and mono listening copies | [sound design](sound-design.md), [waveform editing](waveform-editing.md) |
| Inspection/output | PNG preview, labeled project/solo filmstrips, H.264/AAC MP4, ProRes 4444 MOV (transparent with silent solo), portable `.tsrct` document | [local operation](local-operation.md), [filmstrip review](filmstrip-review.md) |

## Current boundaries

- A document owns one FX composition. Use groups for scenes and Video layers for
  cuts, not `baseVideoTrack`, `fxCompositions`, or caption/audio tracks.
- Import commands package local MP4/MOV/M4V footage and TTF/OTF/TTC font files.
  Installed system fonts are not discovered
  automatically. Audio and images use [media import](media-import.md).
  Import helper-generated files before placing layers.
- The CLI exports MP4 or ProRes 4444 MOV at automatic resolution/frame rate; inspect
  encoded dimensions rather than assuming they equal the canvas. Audio-only export, range export,
  and geometry inspection are not exposed.
- Canvas choices are listed in [local operation](local-operation.md). A bottom
  shape provides a different background; the renderer requires an absent or
  opaque-black document background.
- Native masks/mattes, effects, and supported 3D layer transforms are available
  only as allowed by the installed schema and required local resources. Verify
  person segmentation's model/resource availability and a real matte before
  planning around it. Do not approximate a person with a hand-drawn silhouette.
- JS animation is not HTML/React rendering. 3D layer transforms are not mesh
  import, rigging, simulation, or ray tracing.
- No generative video, avatar service, transcription, dedicated dereverberation
  model, cloud rendering, publishing, or editable Premiere/AE export is bundled.
- Native gain envelopes are timed ducking, not signal-driven compression. Do not
  invent DSP fields. Local derivatives use available FFmpeg and `project import-asset --kind audio`; see [speech cleanup](speech-cleanup.md).
- CLI setup follows [installation](installation.md). Optional audio helpers need
  Python 3.9+ and local FFmpeg/ffprobe. Font import requires a local TTF/OTF/TTC file;
  subsequent rendering uses the document's packaged bytes.

## Look up only what the edit needs

Save the two schemas locally and search for the specific layer, property, effect,
or action definition. Preserve existing fields and animation relationships.
The action schema includes only supported standalone operations; use
checkout/commit for document-level edits and media placement. Use the installed CLI’s schemas and available tools. If the requested feature is unavailable, explain the specific
limit and continue supported work without switching engines silently.
references/cli-version.txt
0.1.0
references/custom-shader.md
# Custom shader effects

Author a new `customShader` when no supported native layer effect expresses the
requested visual operation. It is the general effect primitive underlying many
engine-provided effects.

## Model

```text
CustomShader
├── name             stable upsert label
├── description      agent/human-facing intent
├── wgsl             fragment shader module
├── params[]         scalar uniform inputs
└── textureInputs[]  additional textures bound from layers
```

The owning layer's rendered output is the implicit primary texture.

## Minimal shape

```json
{
  "id": 40,
  "effect": {
    "type": "customShader",
    "name": "exampleTint",
    "description": "Multiplies the layer RGB by an animatable amount.",
    "wgsl": "...complete WGSL module...",
    "params": [
      {
        "name": "amount",
        "description": "Tint multiplier; 0 is unchanged.",
        "min": 0,
        "max": 1,
        "default": 0.5
      }
    ]
  }
}
```

`name` is the upsert key, not the animation address. Animate a parameter using
the enclosing effect instance's stable `id` and the parameter `name`.

## Scalar uniform inputs

Each `params[]` entry declares one scalar `f32` input:

```text
name         shader-local field and animator parameter name
description  explain units and visual meaning
min/max      authoring/interpolation range
default      value used when no animator drives the parameter
```

Parameters are packed into the WGSL `Params` uniform in declaration order.
Keep the JSON order and WGSL struct field order identical. The runtime pads the
uniform payload to the required alignment.

Animate an ordinary effect parameter through the installed action schema's
`setFxPropertyAnimator` effect-property target, using the owning layer ID,
effect ID, and parameter name. Use `layerTimeJsCode` for a script; do not copy an
internal `PropertyTarget` or `animator.code` object into an action. Read
[motion](motion.md) and `tsrct project schema` for the exact wire shape.
A shader that only needs a clock can declare the reserved `animationTime`
parameter instead, subject to support in the installed renderer.

## Parameter naming drives editor controls

Editors classify parameters by **name** and pick an After Effects-style
control. Prefer the explicit control suffix — a `.suffix` on the parameter
name declares the control outright, with no heuristics involved:

```text
<base>.colorR/.colorG/.colorB  color picker — declare each channel 0..1
                               (read as normalized RGB); optional
                               <base>.colorA adds alpha
<base>.x + <base>.y [+ .z]     2- or 3-value row of paired inputs
<base>.angle / .direction /    angle dial in degrees — all six render the
  .rotation / .hue / .phase /  same control; the word records the
  .evolution                   parameter's meaning
<base>.toggle                  on/off switch writing 0 / 1
<base>.slider                  slider paired with a numeric input
<base>.number                  numeric input alone
anything else                  slider + numeric input when the declared
                               range is sweepable, else numeric input alone
```

The suffix never reaches the user: labels strip it (`wet.toggle` reads as
"Wet"; `tint.colorR` as "Tint R" inside the color group; `spin.rotation`
as "Spin"). Choosing among the angle words documents intent for the next
reader of the document — it does not change the control today. WGSL is unaffected —
struct field names are yours and packing is positional; only the JSON `name`
carries the suffix. Grouped suffixes need their sibling set complete
(`.colorR/.colorG/.colorB`; `.x` + `.y`); an incomplete group falls back to
numeric inputs.

The declared range drives the angle dial's write semantics. Declare exactly
`0`/`360` for a modular direction (the shader consumes it through
`sin`/`cos`, so a full turn returns to the same result): the dial and its
numeric input wrap — typing 450 lands on 90. Any other declared range —
including deliberately narrow ones like -45..45 — means the parameter is
consumed linearly (a signed twirl-style amount where -120 and 240 render
differently): both inputs honor the declared `min`/`max` and clamp, so
negative, extended, and narrow values stay exactly as authored. Omitting
`min`/`max` on an angle-suffixed parameter leaves the engine defaults of
exactly 0/1, which editors read as "no range declared" and treat as the
full-turn wrapping encoding — but declare `0`/`360` explicitly so the
intent survives review. Editors only infer a dial for un-suffixed names when
`max` ≤ 360 — multi-turn spans like a 0..720 shutter duration stay numeric
inputs.

Color channels are edited through an 8-bit hex swatch, so a swatch edit
quantizes each channel to 1/255 steps; the per-channel numeric rows keep
full float precision.

The declared range is what earns a plain scalar its slider: bounds must be
finite, `min` below `max`, and both within ±1000. The editor owns that
threshold — `SLIDER_BOUND_LIMIT` in its param classifier — because
`EffectParam` carries no control hint for the engine to be authoritative with;
this document tracks the editor, not the reverse. Past that the bound reads as a
safety cap on a coordinate or technical space (a ±5000-pixel offset, a 0..100000
iteration count) rather than an endpoint anyone sweeps toward, so those get the
numeric input alone. The pair exists because each control reaches what the other
cannot: the slider sweeps the whole range in one control width, the input hits
the exact values the slider's pixel resolution skips.

`.slider` and `.number` force that decision either way — a slider on a range
wider than the guard allows, or a plain input for a parameter meant to be typed
exactly. Unlike the other suffixes they name a widget rather than a meaning,
because a fraction, a level in pixels, and a sample count all want the same
control. Use them only to override; a well-declared range needs neither.

## Reserved parameter names

One parameter name is reserved by the engine: `animationTime`. A parameter
declared with that exact name is not authored, animated, or edited — the engine
overwrites its uniform slot every frame with the owning layer's animation clock
as an `f32` in seconds. It is the same value a `jsScript` effect-parameter
animator on that layer reads as `input.time.seconds`: the local
composition/group clock, rebased to the nearest enclosing Group's in-point and
clamped to that group's duration, with a Group's own effects using the group's
own clock. With no enclosing Group it is the composition playhead. Declare it and
the shader animates with no animator anywhere in the document.

The reservation is total:

- No editor control appears for it. It is filtered out of the effect's parameter
  schema, so agents do not see it either.
- Writes are rejected: effect-parameter update actions refuse the name, and an
  `effectProperty` animator cannot bind to it.
- `min`, `max`, and `default` are documentation only. The engine overwrites the
  packed slot every frame whatever they say — declare a truthful range anyway so
  the intended time domain survives review.
- The parameter must stay declared in `params[]`. Uniform packing is positional,
  so dropping it or reordering around it shifts every later parameter's slot.

```json
{
  "type": "customShader",
  "name": "examplePulse",
  "description": "Pulses the layer brightness on the layer's own clock.",
  "wgsl": "...complete WGSL module...",
  "params": [
    {
      "name": "animationTime",
      "description": "Engine-driven layer animation clock, in seconds.",
      "min": 0,
      "max": 3600,
      "default": 0
    },
    {
      "name": "rate",
      "description": "Pulses per second.",
      "min": 0.1,
      "max": 10,
      "default": 2
    }
  ]
}
```

In WGSL the slot is an ordinary `f32` uniform in the positional `Params` struct —
nothing about reading it differs from any other parameter:

```wgsl
struct Params {
  animationTime: f32,
  rate: f32,
  _pad0: f32,
  _pad1: f32,
};
```

### Shaping time in WGSL

Shape time inside the shader, from `animationTime` plus scalar knob params —
not with an animator. Each pattern is one expression; expose the knob as a
normal parameter so it gets a control:

| Want | WGSL (`let t = params.animationTime;`) | Knob param |
|---|---|---|
| Speed / rate | `t * params.rate` | `rate` (cycles/sec) |
| Offset / delay | `max(t - params.delay, 0.0)` | `delay` (sec) |
| Loop | `fract(t / params.period)` | `period` (sec) |
| Ping-pong | `abs(fract(t / params.period) * 2.0 - 1.0)` | `period` (sec) |
| Hold / stop-motion | `floor(t / params.holdTime)` | `holdTime` (sec/step) |

Use a `jsScript` animator on a normal parameter only for what a pure function
of the clock cannot express: values keyframed by hand, or values driven by
another declared property dependency; there is no automatic audio-analysis input.

## Complete single-texture WGSL template

The owning layer is the only texture when `textureInputs` is empty. This module
is complete and uses the required vertex/fragment entry points, binding layout,
uniform padding, and premultiplied-alpha output contract:

```wgsl
struct VertexInput {
  @location(0) position: vec2<f32>,
  @location(1) tex_coords: vec2<f32>,
  @location(2) color: vec4<f32>,
};

struct VertexOutput {
  @builtin(position) clip_position: vec4<f32>,
  @location(0) tex_coords: vec2<f32>,
};

@vertex
fn vs_main(input: VertexInput) -> VertexOutput {
  var out: VertexOutput;
  out.clip_position = vec4<f32>(input.position, 0.0, 1.0);
  out.tex_coords = input.tex_coords;
  return out;
}

struct Params {
  amount: f32,
  _pad0: f32,
  _pad1: f32,
  _pad2: f32,
};

@group(0) @binding(0) var t_texture: texture_2d<f32>;
@group(0) @binding(1) var s_sampler: sampler;
@group(0) @binding(2) var<uniform> params: Params;

fn author_frag(uv: vec2<f32>) -> vec4<f32> {
  let content = textureSample(t_texture, s_sampler, uv);
  let grayscale = vec3<f32>(dot(content.rgb, vec3<f32>(0.299, 0.587, 0.114)));
  return vec4<f32>(mix(content.rgb, grayscale, params.amount), content.a);
}

@fragment
fn fs_main(input: VertexOutput) -> @location(0) vec4<f32> {
  let raw = author_frag(input.tex_coords);
  // The renderer uses premultiplied-alpha blending. Transparent pixels must
  // carry zero RGB, and every RGB channel must stay at or below alpha.
  return vec4<f32>(
    clamp(raw.rgb, vec3<f32>(0.0), vec3<f32>(raw.a)),
    raw.a,
  );
}
```

Each scalar parameter adds one `f32` field in `params[]` order. Pad the WGSL
`Params` field count to a multiple of four `f32` values (`_pad0`, `_pad1`, …).

## Additional texture inputs

`textureInputs[]` declares named texture slots supplied by other layers:

```json
"textureInputs": [
  {
    "name": "map",
    "description": "Grayscale displacement map.",
    "sourceLayerId": 12
  }
]
```

- `sourceLayerId` references a layer in the same composition.
- Slots resolve in declared order.
- Every required slot must be bound for the shader to run.
- At most eight additional texture inputs are supported.
- A successfully bound source layer is consumed as shader input and is not
  separately painted in the normal layer stack.
- Unknown references, cycles, unbound slots, or excess inputs make the shader
  fall back to a no-op/pass-through behavior.

### Multi-texture binding layout

For `N` additional inputs, `K = N + 1` total textures:

```text
binding 0       owning layer content texture
binding 1..N    textureInputs[] in declaration order
binding K       shared sampler
binding K + 1   Params uniform
```

For example, one additional displacement map uses content at binding 0, the map
at binding 1, the sampler at binding 2, and `Params` at binding 3:

```wgsl
@group(0) @binding(0) var t_content: texture_2d<f32>;
@group(0) @binding(1) var t_map: texture_2d<f32>;
@group(0) @binding(2) var s_sampler: sampler;
@group(0) @binding(3) var<uniform> params: Params;
```

The `textureInputs[]` names describe logical slots; WGSL variable names are
chosen by the shader author. Declaration order, not the logical name, determines
the binding number.

## Authoring checklist

1. Inspect the installed document schema's layer-effect definitions for an
   existing effect before authoring a `customShader`.
2. Give the shader a stable, descriptive `name` and explain its intent.
3. Describe every parameter, including units and useful range, and declare
   its control with an explicit name suffix (`.colorR/.colorG/.colorB`,
   `.x`/`.y`, `.angle`, `.toggle`) so it gets an intuitive control. Declare a
   real `min`/`max` on every plain scalar — that is what earns it a slider.
4. Keep parameter declaration order aligned with the WGSL `Params` struct.
5. Describe every texture input and bind it to a valid source layer.
6. Preserve premultiplied alpha: output RGB must not exceed output alpha.
7. Keep sampling and loops bounded for render-time performance.
8. Animate knobs through `effectId + paramName`, not by rewriting WGSL per frame.
   When the shader needs a clock, declare the reserved `animationTime` parameter
   rather than an animator.
9. Check `tsrct project schema --document`, save through `project commit` or
   a supported action batch, then render with `tsrct preview` so the renderer
   validates the WGSL and bindings. Inspect the actual effect; a pass-through
   fallback is not a successful shader.

## Prefer self-describing shaders

Descriptions travel with the document and become context for future agents.
Write what an input means, not merely its type:

```text
Bad:  "map texture"
Good: "Grayscale displacement map; lighter pixels shift samples farther right."
```
references/editorial-decisions.md
# Choose the treatment from the job

When music is already selected, establish its phrase/beat map before planning visual timing. When tightening speech, inspect waveforms before choosing trims and verify the rendered joins afterward. Follow [waveform editing](waveform-editing.md) for the actual local views and source/edit timing checks.

## Three modes, one video

| Mode | Use it when | What would make it wrong |
|---|---|---|
| Footage alone | The action, expression, product behavior, or spoken thought carries the point. | Decorating a strong moment distracts from the evidence. |
| Footage with an overlay | The viewer needs a name, emphasis, measurement, comparison, pointer, or context while seeing the real action. | The graphic repeats speech, hides the subject, competes with captions, or becomes a full-screen card by accident. |
| Full motion-graphics scene | The idea is abstract, a relationship needs explaining, there is no useful footage, or a chapter/brand moment needs a deliberate visual reset. | Interrupting emotional or product evidence with a generic title card, or turning the edit into consecutive slides. |

Choose the amount from the material, not a fixed footage-to-graphics ratio. A dialogue edit may need almost no graphics. A technical explainer can need several full scenes. A motion-only brief may use none of the footage. Preserve those distinctions.

## Source selection comes before embellishment

Inspect enough of each clip to identify usable in/out points, action, camera movement, eyelines, and audio. Keep a compact source log with filename, source time range, and why the moment matters. A contact sheet alone cannot establish speech meaning or whether an action completes cleanly.

For speech: preserve complete thoughts and natural breaths. Remove redundancy with care; an aggressive silence cut can sound fragmented. Use an L-cut or J-cut only when it improves continuity and keep source audio mapped explicitly. Use B-roll to support what is being said, not just to hide every edit.

For action: cut on motivated movement or changes in information. Leave enough pre-action and post-action handles to read the motion. A speed ramp is purposeful emphasis; do not accelerate a shot merely to fit an arbitrary runtime.

For products: show actual behavior clearly before calling it a feature. Exact UI text can come from supplied screens/assets. Avoid invented product screens, metrics, quotes, or claims.

For ads, use [ad hooks](ad-hooks.md) to select the opening idea and align it to the music/action before expanding the full edit plan. Choose the source evidence and viewer takeaway first, then the motion and sound treatment.

## Talking-head eyeline continuity

When cutting between shots of the same speaker, align the speaker's eyes across the cut for a natural transition. Compare their eye height and horizontal screen position in the **final composition**, after crops, scaling, and aspect-ratio changes. Use the midpoint between visible eyes as a practical reference rather than the top of the head, which changes with head tilt and framing.

Inspect the outgoing and incoming frames plus a short interval on either side; a blink or passing nod should not determine the framing of an entire shot. Apply modest per-shot position/crop and uniform scale adjustments to keep the eyeline at a consistent screen height and avoid an accidental sideways jump. Preserve useful headroom, looking room, facial proportions, and source resolution. For an intentional punch-in, allow the face size to change while keeping the eyes near the same screen position. A close-up need not match the wide shot's head size.

Keep natural head movement; do not introduce jitter by chasing the eyes frame by frame. Screen position and gaze direction are separate: reframing cannot correct a change in where the speaker is looking. Respect deliberate camera-angle/framing changes and interview looking direction; do not mirror or alter the speaker's eyes to manufacture continuity. If matching requires an awkward crop or excessive enlargement, choose a better source boundary or shot while preserving complete words and audio sync.

Follow [filmstrip review](filmstrip-review.md) to compare the last outgoing and first incoming frames at the same display size, then play the join at normal speed. Recheck adjacent cuts after changing a shot's framing, including captions and overlays that may now collide with the face. Record any intentional framing change that remains.

## A compact plan

For a new edit a table like this is sufficient; keep it proportional to complexity:

| Edit time | Source moment | Viewer takeaway | Treatment | Sound |
|---|---|---|---|---|
| 0–3s | Close-up where the mechanism engages | See the result immediately | Footage; no title obstruction | Keep the actual click |
| 3–7s | Wider view / explanation | Understand how it works | Small tracked-in-space callout if supported, otherwise a stable pointer | Dialogue forward |
| 7–10s | No footage explains the relationship | Understand the two-part process | Full native diagram with a readable hold | Continue the music bed |
| 10–15s | Return to the result | Connect explanation to proof | Footage; restrained end label | Resolve naturally |

These numbers are illustrative, not a template or required pacing.

For each beat, consider the sound roles in [sound design](sound-design.md): primary voice, music, a small visual accent, a structural cue, and physical texture. Mark only those that earn a place, plus intentional silence. Plan the timing of a sound’s perceptual peak against the action, then audition it.

## Rhythm and continuity

- Identify phrases, downbeats, impacts, and quiet intervals in supplied music. Let those inform structure; avoid cutting on every beat.
- Establish one main transition language tied to the content: a match cut, shared shape, directional movement, or sound bridge. Cut directly where a transition adds nothing.
- Reuse a few design decisions—type hierarchy, grid, palette, motion character—so scenes feel related. Repetition of a design system is useful; repetition of the same scene arrangement is usually not.
- Build toward a specific payoff. Keep some scenes quieter so the important moment has contrast.
- A requested duration is a delivery constraint. Achieve it through selection and timing, not a globally sped-up finished video or unreadably short text.

For a revision, change only the necessary source ranges/layers/animators, retain the existing language, and inspect neighboring beats for continuity.
references/filmstrip-review.md
# Mid-edit review with Tesseract filmstrips

Tesseract can render sampled project frames into a labeled PNG grid without encoding a full MP4. Use this during editing to catch problems before building more work on top of them.

## The review loop

After the rough assembly and after a meaningful visual change (a new graphic, text layout, crop, timing, or transition):

1. Save the current `.tsrct` revision through `project commit` or `project apply`. The strip must come from that revision, not a stale document or the source footage.
2. Render a filmstrip covering the changed beat and its neighboring cuts. For a rough edit, first make a sparse overview covering the sequence.
3. **Open the PNG with the host's image-view tool and inspect the actual pixels.** A successful command or an image path alone does not count as review. In Codex, use the available local image viewer, such as `view_image`.
4. Identify the timestamp and observable problem: text collision, premature exit, blank frame, wrong crop, inconsistent scale, missing asset, or distracting repetition. Correct the native project, save, and render a fresh strip of the affected range. Compare before/after with the same sample times.
5. Keep a short note identifying the inspected project revision, strip path, times, issues fixed, and any remaining concern. This can live in the edit notes; no separate formal report is required.

Do not run a new strip after every scalar tweak. Review coherent changes while they are still cheap to correct. Finish by playing a short preview or the final edit with audio; a filmstrip cannot establish sound quality, smooth motion, or lip sync.

## Choose useful samples

- **Assembly overview:** roughly every 1–2 seconds, plus important scene boundaries. Split long edits into readable strips rather than making one giant sheet.
- **A changed beat:** entrance, mid-motion, readable hold, and exit, with context from the preceding/following shot. About 100–250 ms spacing is useful for a short reveal; use explicit times when the important states are known.
- **A suspect cut:** sample the frame just before and just after the cut, as well as the cut time. At 30 fps, a frame is about 33.3 ms; for a cut at 2.0 s, 1967, 2000, and 2033 ms are useful requests. Inspect the resolved labels and source cadence; a rounded millisecond request is not a promise of exact source-frame indexing.
- **Talking-head cuts:** compare outgoing/incoming frames and a few surrounding samples at the same tile size for eye height, horizontal position, and headroom. Inspect the final crop, including punch-ins; apply [eyeline continuity](editorial-decisions.md#talking-head-eyeline-continuity), then re-render and play the join to check for an unnatural face jump.
- **Small text or matte detail:** request a larger tile, a native crop, or a full-resolution `preview`. Thumbnail text is not enough evidence of final-size legibility.

Uniform sampling can miss a one-frame flash. It also cannot prove a complete absence of overlap between samples. Use denser targeted samples or playback when the behavior is uncertain.

## Tesseract CLI

A readable landscape overview:

```sh
tsrct filmstrip --project /absolute/path/edit/project.tsrct --start-ms 0 --duration-ms 8000 --interval-ms 1000 --output /absolute/path/edit/Previews/Filmstrip-v2.png
```

A closer review around a cut:

```sh
tsrct filmstrip --project /absolute/path/edit/project.tsrct --timestamps-ms 1700,1900,1967,2000,2033,2100,2300,2600 --output /absolute/path/edit/.tesseract-work/checks/cut-v2.png
```

Keep ranges inside the actual document duration. Range sampling is half-open;
explicit timestamps preserve input order. Do not sample exactly at the edit end
and mistake an inactive end state for a bad final frame. About 8–16 samples are
usually easier to inspect than a giant sheet.

Cells default to 240×135 with five columns and ordinal/time labels. Use
`--tile-width 135 --tile-height 240` for portrait; match the project aspect for
square/4:5 work. Increase tile size or use `tsrct preview --time <seconds>`
when details become too small. Choose fresh output filenames for comparisons.
Stdout JSON reports `outputPath`, `frameCount`, dimensions, `resolvedSampling`,
and `resolvedFraming`. Use resolved times and labels for cut-level diagnosis.

## Fonts must match the final render

Package the intended local file with `project import-font --file` before authoring
text. Preview, filmstrip, and export use that document's packaged fonts; there
is no separate runtime-font mapping flag. Inspect the glyphs, line breaks,
and overflow. If import or font resolution fails, report the exact error and
follow [local operation](local-operation.md); do not silently substitute a font.

## Targeted inspection

Use an actual composition/group ID from inspection:

```sh
tsrct filmstrip --project /absolute/path/edit/project.tsrct --fx-solo main:3 --max-frames 8 --content-bounds --tile-width 320 --tile-height 180 --items-per-row 4 --output /absolute/path/edit/.tesseract-work/checks/graphic-v2.png
```

This samples the selected layer/group's full active window. Omit
`--content-bounds` to preserve its canvas placement. Use one sampling mode at a
time. A solo strip is for diagnosis: also inspect the graphic in the complete
composition before calling it finished. This CLI does not expose arbitrary crop
or target-offset flags.
references/fx-authoring.md
# FX authoring

Use `project inspect` for existing content and `project schema` for the installed
CLI's action contract, or `project schema --document` for editable JSON.
The footage example edits document JSON; the other examples are action batches.

## Structure, placement, and resources

- The document owns one composition starting at time zero. Its root `duration`
  is seconds; layer `activeRange` and `sourceRange` values are milliseconds and
  immediate-parent-local. Keep document and ancestor windows long enough.
- Layer IDs are integers unique within a composition. Choose unused IDs after
  inspection; preserve IDs when editing. Names describe visible roles.
- Sibling index zero is frontmost. Create actions append behind siblings when
  `insertIndex` is omitted; use `insertIndex: 0` for a foreground insertion.
- `position` is in parent coordinates; `anchorPoint` is in layer-local coordinates.
  Scale and opacity use percentages (100 is unchanged/opaque); RGBA uses 0–1.
  Read the project's canvas dimensions rather than treating source pixels as
  composition coordinates. Parent transforms also affect children.
- `sourceRange` selects media time; `activeRange` places it in the composition.
  Moving content and trimming its source are different operations.
- Import local footage with `project import-video` before referencing its asset ID.
  An arbitrary path or asset ID in an action does not import its bytes.
  Import audio/images through [media import](media-import.md); package fonts as below.

## Local footage

Package supplied footage first; the command saves the video bytes inside the
project without modifying the source file:

```sh
tsrct project import-video --project project.tsrct \
  --file /path/to/clip.mp4 --asset-id footage-1
```

Use a new ID containing letters, numbers, hyphens, or underscores. Existing IDs
are rejected to avoid replacing another layer's media. The JSON response reports
`assetId`, `durationMs`, `width`, and `height` (display orientation). MP4, MOV, and
M4V containers are accepted; actual decoding uses the engine's supported codecs.
No external metadata tool is required. Import adds the resource; the following
editable document JSON places it on the timeline. Run `project checkout`, append
the layer below to `composition.layers`, set document `duration` to `3` seconds,
then run `project commit --project project.tsrct --file .tesseract-work/editable.json`. Preserve
other layers; append footage behind overlays because index zero is frontmost.

For a three-second clip, for example (adapt duration, IDs, canvas, and transform
to the returned metadata and the user's requested framing):

```json
[
  {"type":"Video","id":1,"name":"Main footage",
   "activeRange":{"start":0,"duration":3000},"sourceRange":{"start":0,"duration":3000},
   "sourceIntrinsicDuration":3000,"volume":1.0,
   "transform":{"anchorPoint":[0,0],"position":[0,0],"scale":[100,100],"rotation":0,"opacity":100},
   "source":{"assetId":"footage-1","fit":"contain"}}
]
```

Import each distinct file once; multiple layers can reference the same asset.
Use the actual clip duration for `sourceIntrinsicDuration`, and select the wanted
section with `sourceRange`. Preserve footage audio unless the user asks to mute
or replace it. Set `volume: 1.0` on new Video layers to enable embedded audio
at unity gain; an omitted/null volume leaves it disabled. Preserve existing
gain and animators on revisions, and verify the exported audio.
Inspect a preview to verify crop, orientation, and composition placement before
adding overlays or exporting. If decoding fails, report the error and source format.

## Rectangle overlay

In the default `main` composition, this adds a blue card. Adapt the inspected
composition ID, unused layer IDs, and geometry to the existing document.

```json
[
  {"type":"createFxRectLayer","compositionId":"main","layerId":1,"name":"Card","insertIndex":0,
   "activeRange":{"start":0,"duration":3000},
   "transform":{"anchorPoint":[0,0],"position":[80,80],"scale":[100,100],"rotation":0,"opacity":100},
   "rect":{"size":[400,160],"fillColor":[0.05,0.25,0.8,1]}}
]
```

## Group related parts

After the previous example, add an accent in front and group the two siblings.
The identity group transform preserves their placement. Animate the group to
move the callout together; animate a child for a detail within it.

```json
[
  {"type":"createFxRectLayer","compositionId":"main","layerId":2,"name":"Accent","insertIndex":0,
   "activeRange":{"start":0,"duration":3000},
   "transform":{"anchorPoint":[0,0],"position":[80,80],"scale":[100,100],"rotation":0,"opacity":100},
   "rect":{"size":[8,160],"fillColor":[1,0.7,0.1,1]}},
  {"type":"groupFxCompositionLayers","compositionId":"main","layerIds":[2,1],"groupLayerId":3,"name":"Callout group",
   "transform":{"anchorPoint":[0,0],"position":[0,0],"scale":[100,100],"rotation":0,"opacity":100}}
]
```

Grouping requires contiguous siblings with the same parent. Use the grouping
operation rather than manually reparenting children and rewriting their timing.

## Text with an available font

Choose typography for the brief before acquiring files. Honor supplied brand
fonts; otherwise choose the character, width, weight, and hierarchy that suit
the actual headline and supporting copy. The example's Inter Bold is a wire-format
example, not a recommended design default. Availability in an OS folder, cache,
or test fixture is not a reason to choose a face.

Use supplied font files or obtain the selected TTF/OTF/TTC from an official
source such as Google Fonts when downloads are allowed. Keep a recognizable
filename and the source/license with the working files. If the requested font
cannot be obtained, explain the gap before substituting it.

Import the selected files into the document before authoring text:

```sh
tsrct project import-font --project project.tsrct --file /path/to/Inter-Bold.ttf
```

Use the returned `fontFamily` and `fontStyle`, not names guessed from the filename.
For a collection, choose the intended entry from `faces`. The filename is a
readable asset label; parsed metadata identifies the face, and the asset ID
identifies the file's bytes. Import only the faces needed for the chosen design.
Repeat import is a no-op. Preview and export use the packaged bytes after the
document moves to another machine.

Before building text-led animation, preview the actual words at their intended
size and placement. Check character, hierarchy, line breaks, and glyph coverage;
if the choice is uncertain, compare a small number of candidates in a scratch
project before importing the selection into the deliverable. A successful import
proves the file is usable, not that its typography fits the brief.

This adds box text to group 3 above using the imported Inter Bold face.

```json
[
  {"type":"createFxTextLayer","compositionId":"main","parentLayerId":3,"insertIndex":0,"layerId":4,"name":"Callout label",
   "activeRange":{"start":0,"duration":3000},
   "transform":{"anchorPoint":[0,0],"position":[112,112],"scale":[100,100],"rotation":0,"opacity":100},
   "sourceText":{"text":"A useful detail","fontFamily":"Inter","fontStyle":"Bold","fontSize":32,
     "fillColor":[1,1,1,1],"justification":"left","boxText":true,"boxPosition":[0,0],"boxSize":[336,96]}}
]
```

Box text wraps but does not automatically shrink or clip overflow. Point text's
local origin is its baseline start, not its visual center. Inspect a rendered
frame for font metrics, line breaks, and overflow before animating the layout.


## Bottom captions

Keep spoken captions editable as FX text layers, with one short phrase per
`activeRange`. Obtain the actual words and timings from a supplied transcript,
reviewed audio, or an available transcription tool; the CLI does not transcribe.
If none is available, ask for the spoken text rather than inventing dialogue.
Map source timestamps through the footage's `sourceRange` and `activeRange`.

Import the selected legible font file first. On a 1080×1920 portrait canvas, a useful
starting point is centered box text at `[90, 1550]`, box size `[900, 220]`,
font size `64`, and white fill. Adapt to the inspected canvas and shot. Keep
phrases to one or two lines and use a dark rounded rectangle behind them when
footage contrast needs it. Put both layers above the footage with matching
active ranges, and keep the text ahead of its background in sibling order.

Review phrase boundaries and the longest caption in the composed filmstrip.
Verify the text stays inside the frame, clears the face and bottom edge, and
matches the audible words. Preserve the video's embedded audio gain.
references/installation.md
# CLI installation

The skills and plugin contain instructions, not the CLI. First look for the installed
command (`~/Library/Application Support/Tesseract/bin/tsrct` on macOS,
`%LOCALAPPDATA%\Tesseract\bin\tsrct.cmd` on Windows), then `tsrct` on PATH,
and run `--version`.

Read the required CLI version from `cli-version.txt` beside the installed
installation guide. If the pin is missing, report an incomplete release package
and stop installation. Check the executable's reported version against the pin;
do not substitute the latest release or an internal build.

The CLI supports macOS and 64-bit Windows. Pick the bundle before installing:

- macOS: `arm64` for Apple Silicon and `x86_64` for Intel. Check `uname -s` and
  `uname -m`. `sysctl -in sysctl.proc_translated` returning 1 means Apple
  Silicon under Rosetta; use the `arm64` bundle.
- Windows: `windows-x86_64` for 64-bit Windows 10 or later. Check
  `$env:PROCESSOR_ARCHITEW6432` when set (a 32-bit PowerShell on 64-bit Windows),
  otherwise `$env:PROCESSOR_ARCHITECTURE`; either must be `AMD64`. Nothing
  needs to be preinstalled; the bundle carries its runtime libraries.

On any other host, report that this host is unsupported and stop installation.

For a missing or mismatched pinned CLI, download the required bundle, then
verify and install it as described below.

## Public CLI downloads

Read the exact required version from `cli-version.txt` beside the installation guide.
If the pin is missing, report an incomplete release package and stop installation.
Do not substitute the latest release or an internal build.

Download these two files from
`https://github.com/mirage-hq/tesseract/releases/download/v<VERSION>/`:

- `tesseract-<VERSION>-<PLATFORM>-<ARCH>.zip`
- `tesseract-<VERSION>-<PLATFORM>-<ARCH>.zip.sha256`

Use `darwin-arm64`, `darwin-x86_64`, or `windows-x86_64` for
`<PLATFORM>-<ARCH>`, following the host checks in [installation](#cli-installation).
Use an available downloader (for example, `curl -fL` on macOS or
`Invoke-WebRequest` on Windows) or a browser. GitHub login and GitHub CLI are
not required. Save both files in a fresh download directory.

If network permissions block the download, use the agent's normal permission
flow. If the release or platform asset is unavailable, report the exact URL and
error and stop installation; do not request private repository access or guess
another version. The release page is
`https://github.com/mirage-hq/tesseract/releases/tag/v<VERSION>`.

Return to [installation](#cli-installation) to verify the checksum, extract,
install, and check the executable's version. Downloaded archives are not installed
until those checks pass.

## Verify and install

1. macOS: from the download directory run `shasum -a 256 -c <archive>.zip.sha256`,
   extract with `ditto -x -k <archive>.zip <destination>`, then run
   `bash <extracted-CLI-folder>/install.sh`.
   Windows: compare `(Get-FileHash <archive>.zip -Algorithm SHA256).Hash` with the
   `.sha256` file, extract with `Expand-Archive <archive>.zip <destination>`, then run
   `powershell -ExecutionPolicy Bypass -File <extracted-CLI-folder>\install.ps1`.
2. Run the installed command with `--version` using its quoted absolute path
   (`"$HOME/Library/Application Support/Tesseract/bin/tsrct"` or
   `"$env:LOCALAPPDATA\Tesseract\bin\tsrct.cmd"`) and use that path for
   subsequent commands.

The installers check platform, architecture, minimum OS version, and bundle
checksums before switching the installed command. No sudo, administrator rights,
Rust, Homebrew, or Node is needed. Mac releases are Developer ID signed and
Apple notarized; first launch may require
internet access for notarization checks, and macOS may still block launch
depending on device policy. The Windows bundles are unsigned, so SmartScreen may
warn. If blocked, let the user review the warning and follow their
organization's approval process. Do not bypass Gatekeeper or
SmartScreen or delete quarantine attributes. Report the exact error if a download,
installation, or launch fails.
references/local-operation.md
# Running Tesseract locally

Resolve this skill from the directory containing its `SKILL.md`. In the sound
examples, `SKILL` means that absolute directory, whether installed alone or in a
plugin. Quote paths containing spaces. The optional sound/waveform tools use
Python 3.9+ and local FFmpeg/ffprobe; they do not install a renderer. If a tool is
missing, report the specific unavailable check and use the host's available
media tools where they can establish the same result.

Tesseract uses `tsrct` to edit and render portable `.tsrct` projects. Follow
[CLI installation](installation.md) to locate the executable,
check its version, or install the matching CLI release. Use the resolved
executable path in the commands below; it does not have to be on PATH.

## Scope

For creating or restructuring layers, read [FX authoring](fx-authoring.md).
For adding or revising animation, read [Motion](motion.md).

Use editable FX compositions, text, shapes, media layers, effects, keyframes,
and procedural scripts. Use existing local media. This CLI has no asset
generation, transcription, dubbing, or cloud job tools.

## Project lifecycle

Choose the intended project before writing. Reuse an existing project only when
the conversation, supplied paths, or project notes establish continuation; a
matching name alone is insufficient. Inspect that destination before modifying it.
Ask only when intent to modify existing work remains ambiguous.

For new work, claim a fresh project root with an atomic operation that fails if
the path already exists, such as `mkdir "$project_root"` without `-p`. Check that
it succeeded before writing any project files or creating subfolders. If the path
exists, leave it untouched and choose another name; do not merge new work into it.
`mkdir -p` is appropriate for subfolders only after the project root has been
claimed or established as the intended existing project. A successful `mkdir -p`
or an earlier existence check is not proof that this task created the directory.

Use one persistent project root for authoring, review, revisions, and handoff.
Keep retained working files in subfolders within that root, not in a separate
workspace left behind at delivery. Hand back the same root. A separate client
package is an additional export when requested; relocating the project means
moving the retained workspace together and checking its references.

For example (names are flexible; create only useful folders):

```text
Project/
  Project.tsrct     current editable document with packaged media and fonts
  Project.mp4        matching current video
  Previews/          useful filmstrips, poster frames, or comparisons
  Versions/          meaningful earlier drafts, when needed
  .tesseract-work/      retained notes, scripts, JSON, and diagnostics
```

Keep a readable filmstrip of the delivered revision visible, even if it began as
an internal check. Retain other artifacts according to their value for reviewing,
reusing, or continuing the work. Schema dumps and temporary debugging renders
belong in the working subfolder (or a version-matched schema cache). The `.tsrct`
document is self-contained; loose copies of embedded assets are not sidecars.
Preserve original user media and retain external sources when independently useful.

Save meaningful checkpoints before substantial revisions and keep enough notes
to resume: intent, decisions, current revision, actual checks, and unresolved issues.
Associate exports and previews with their source revision; avoid saving every
small tweak. Tidy disposable intermediates without deleting useful history.
These are workspace conventions, not file-format requirements or mandatory files.

## Before editing

Run `tsrct --version`, `tsrct project --help`, and `tsrct export --help`.
If `import-asset` or `--codec` is missing, follow [installation](installation.md).
The examples below run from the chosen project root; adapt filenames to the task.
Create an empty document for new work, or inspect the intended existing document:

```sh
tsrct project create --project project.tsrct
```

Creation makes one empty composition (`main`) on a 1080×1920 canvas, lasting
three seconds. Inspect its ID before adding content. Change canvas, duration,
or composition name by checking out and committing the editable document JSON.
The document `duration` is in seconds; layer ranges and action times are milliseconds.
Native rendering currently supports 1080×1920, 1920×1080, 1080×1080, 1080×1350,
810×1080, and 1350×1080 canvases. Leave `backgroundColor` absent or opaque black;
use a bottom FX shape layer for another background.

Import supplied footage with `project import-video` before adding its video
layer through JSON checkout/commit. Read [local footage](fx-authoring.md#local-footage).
For images, music, sound effects, and processed audio, follow [media import](media-import.md).
Before adding text, choose and obtain the intended font files, then package them
with `project import-font --file`; see [text](fx-authoring.md#text-with-an-available-font).
Handle setup and editing yourself; the user should only describe the result.
For spoken captions, read [bottom captions](fx-authoring.md#bottom-captions).

Inspect before constructing edits:

```sh
mkdir -p .tesseract-work
tsrct project inspect --project project.tsrct --pretty > .tesseract-work/project.json
tsrct project schema > .tesseract-work/project-action.schema.json
tsrct project schema --document > .tesseract-work/document.schema.json
```

The document contains one composition. Use groups for related parts; do not
create extra compositions or base-video, caption, or audio tracks.

Use IDs from `.tesseract-work/project.json`. Read the relevant action definition from the
schema; do not invent fields.

## Edit loop

For shape, text, grouping, and animation edits, write a JSON array of supported
composition-local actions in `.tesseract-work/edits.json`, applied as one atomic batch.
The action schema excludes operations this standalone document rejects:
project metadata, canvas changes, composition lifecycle, and asset-introducing actions.
For those document edits and media layers, use the editable JSON schema:

```sh
tsrct project checkout --project project.tsrct --output .tesseract-work/editable.json
# Edit dimensions, duration, composition, or layers using .tesseract-work/document.schema.json.
tsrct project commit --project project.tsrct --file .tesseract-work/editable.json
```

Preserve existing animation graphs and unknown fields when editing JSON. Commit
validates the document and packaged asset references before atomically saving.
Re-checkout after action batches so a stale JSON file cannot undo those changes.
Use the working examples in the references, adapting IDs, timing, and geometry
to the inspected project.

Apply and review:

```sh
mkdir -p Previews .tesseract-work/checks
tsrct project apply --project project.tsrct --actions .tesseract-work/edits.json
tsrct preview \
  --project project.tsrct \
  --time 1 \
  --output .tesseract-work/checks/layout.png
```

Inspect `.tesseract-work/checks/layout.png` for layout. Review motion across its active range:

```sh
tsrct filmstrip --project project.tsrct \
  --start-ms 0 --duration-ms 3000 --interval-ms 250 --output Previews/Filmstrip.png
```

Open and inspect the generated images. Follow [filmstrip review](filmstrip-review.md)
for sampling, solo diagnostics, and the correction loop.

## Finish

```sh
tsrct export \
  --project project.tsrct \
  --output finished.mp4
```

For transparent motion graphics on macOS, export ProRes 4444:

```sh
tsrct export --project project.tsrct --codec prores4444 --output finished.mov
tsrct export --project project.tsrct --codec prores4444 --fx-solo main:3 --output overlay.mov
```

Full-project export includes the mix and the black canvas background. For alpha,
use solo export of the overlay layer or a group containing the graphic. Solo uses the selected
layer/group's active window, preserves canvas placement, and omits audio. Inspect
its alpha channel and composite over contrasting backgrounds before delivery.

Export uses the renderer's automatic resolution and frame rate; canvas dimensions
do not guarantee matching encoded dimensions. Inspect the resulting MP4 against
the delivery brief. This CLI does not expose resolution or FPS override flags.

Follow [review and delivery](review-and-delivery.md#handoff) for final checks and
presenting the result from this project root.

The `.tsrct` document is the portable editable source. Checked-out JSON becomes
the saved source only through `project commit`; do not edit the ZIP manually.

## Safety and failure behavior

- Never edit `project.tsrct` as text.
- Preserve the input with `--output .tesseract-work/revised.tsrct` when experimenting.
- `project apply` is atomic: a rejected action batch does not publish partial
  project changes.
- Preserve the `.tsrct` document and its packaged resources.
- This release exposes the supported subset of the engine action schema. Use only actions
  required for the user's request.
- For `missing_fonts`, obtain and import the requested files, then retry.
  Follow [font selection and import](fx-authoring.md#text-with-an-available-font);
  names in the error identify the requested faces.
  If import fails, report the actual error. Do not silently replace a requested face.
- FX text style and layout changes go through `project apply`.

## Media inspection and optional audio helpers

Import reports metadata, not what a shot depicts or what its audio means. Inspect
source frames and listen through the host's media tools before choosing cuts.
For local metadata when FFmpeg tools are available:

```sh
ffprobe -v error -show_streams -show_format -of json /absolute/path/footage.mov
```

Use [waveform editing](waveform-editing.md) for source and rendered audio views,
and [audio and timing](audio-and-timing.md) for sound helpers and gain envelopes.
The helpers synthesize WAV accents on demand and inspect audio; they do not
import media or mutate a project. Import their output with `project import-asset --kind audio`,
then place editable Audio layers. Audio-only and arbitrary range export are not
exposed by this CLI. Use the CLI to modify documents; do not edit the ZIP directly.

For a dialogue-only diagnostic, check out the working document and commit a
separate `.tsrct` copy with music/SFX gains and any gain animators muted. Export
that copy to MP4 and inspect its audio directly with the waveform helper. Preserve
the working mix and source files; do not claim unperformed listening checks.
references/media-import.md
# Import images and audio

Import copies bytes into the portable document. It does not create timeline layers.
Use a unique flat asset ID; existing IDs cannot be replaced accidentally.

```sh
tsrct project import-asset --project project.tsrct --file /absolute/path/logo.png --asset-id logo --kind image
tsrct project import-asset --project project.tsrct --file /absolute/path/music.wav --asset-id music --kind audio
```

Accepted image suffixes: PNG, JPG/JPEG, WebP. Audio: WAV, MP3, M4A, AAC, FLAC, OGG.
The underlying decoder must support the file's encoding; import checks the file
and suffix and packages it, but does not decode or return dimensions/duration.
Probe metadata with local media tools, then verify decoding in a preview/export.
Use `project import-video` for footage and `project import-font --file` for local font files.

Check out the document after import. Add an Image or Audio layer using the
installed document schema and the imported asset ID, then commit. Preserve
existing layers and use unused numeric layer IDs. For a 200×100 image:

```json
{"type":"Image","id":1,"name":"Logo","activeRange":{"start":0,"duration":3000},"transform":{"anchorPoint":[100,50],"position":[320,180],"scale":[100,100],"rotation":0,"opacity":100},"source":{"assetId":"logo","fit":"contain"}}
```

Use actual image dimensions for the anchor and choose placement for the canvas.
For Audio layer fields, source timing, gain, and envelopes, read
[audio and timing](audio-and-timing.md). Generated cues and FFmpeg-processed
derivatives follow the same import path. Keep originals and processing recipes;
mute the original dialogue when replacing it to avoid doubled playback.
references/motion-design.md
# Motion design: full scenes and overlays

## Design the motion before the syntax

State one visual idea in concrete terms: a number grows into a comparison; a line reveals the flow through a process; a label follows the subject's direction; typography is revealed by the product's silhouette. If the only idea is “things fade in,” reconsider the design.

Decide the focal element, its relationship to the footage, and how the composition changes over time. A full scene should have composition, progression, and a payoff, not just a centered heading above three rounded boxes.

For motion that opens a paid-social or YouTube ad, read [ad hooks](ad-hooks.md) for concept selection, placement context, music/SFX synchronization, and the early opening review.

## Overlays

- Start with the footage's actual safe area. Account for the face, hands, product, subtitles, and platform cropping. Check placement across the whole interval, not one still.
- Prefer a small label or shape that points to meaningful evidence. A large opaque panel is appropriate only if the brief calls for that treatment.
- Animate in, hold long enough to read, animate out. Entrance and exit should relate to the cut or the subject's movement.
- Use contrast from placement first, then a subtle scrim/stroke/shadow as needed. Do not blanket every shot in a heavy gradient.
- For graphics behind a person, use a person matte when its required local model/resource is available, and confirm its polarity. Verify support on an actual frame first; the skill does not install segmentation models. A person-white matte with **lumaInverted** reveals graphics outside the person; **luma** reveals them inside the person. Some engine recipe prose describes the latter without stating the intended visual relationship—pick and test the polarity for the actual design.
- For a tracked pointer, do not describe it as tracked unless there are actual time-varying positions aligned to observed footage. Static arrows are fine when appropriate.

## Full-frame scenes

Build a hierarchy: focal content, supporting context, environment. Use spatial and temporal relationships to explain meaning. Possibilities include a native vector schematic, dimensional arrangements of 2D layers, split-screen comparisons, mask-driven type reveals, data with accurate values, or footage integrated into a composited scene.

Reserve motion for changes that guide attention. Give the viewer a stable interval to understand the result. Avoid defaulting every scene to parallax, glow, particles, bouncing labels, and a camera push at once.

## Native mechanism map

| Design need | Tesseract mechanism | Look up |
|---|---|---|
| Type hierarchy / editable text | `Text.sourceText` and imported font files | [FX authoring](fx-authoring.md#text-with-an-available-font) and the document schema |
| Per-character reveals | Text Animator and selectors; drive fields with dynamics | [FX authoring](fx-authoring.md#text-with-an-available-font) and [motion](motion.md) |
| Coordinated movement | `Group` with child layers and stable parent IDs | [FX authoring](fx-authoring.md) |
| Drawn paths / linework | `Shape` paths, fills, strokes; animate supported path fields | the document schema's Shape definitions |
| A moving reveal | Track matte or path mask using a real layer reference | the document schema's mask and track-matte definitions |
| Treatment across several layers | Group effects or an Adjustment layer above the affected siblings | the document schema's effect definitions |
| Soft lighting, blur, grade | Ordered typed layer effects | the document schema's effect definitions |
| 3D-looking layer arrangement | Supported 3D transform properties and depth ordering | the document schema and [motion](motion.md) |
| Custom pixel treatment | Bounded premultiplied-alpha WGSL custom shader | [custom shader contract](custom-shader.md) |

This is not a Blender runtime. Distinguish transforming 2D layers in 3D from general meshes, physics, or a full 3D scene.

## Timing and motion character

Tesseract FX uses editable keyframe actions and `AnimationGraph` expressions, not CSS or AE keyframe-array JSON. Read [motion](motion.md) for the installed action workflow. Use `layerTimeJsCode` and the owning layer's clock. An expression returns its value; it is not a general browser script.

For a simple reveal, 200–450ms is a useful starting range, adjusted to the scale of movement and edit. Give secondary details a small deliberate offset. Choose easing based on mass and purpose; a restrained ease-out usually reads more deliberately than an automatic spring. Overshoot can fit playful material but should be intentional.

Use standard motion blur only where it improves fast transforms; enable both composition and layer switches. Long tails of blur can erase text or mask a timing error. Render the actual movement before committing to the treatment.

## Brand application

Use supplied fonts, palette, logo geometry, and references. Keep exact logo art separate from generated scenery. This plugin's Mirage icon is not a default asset for client work. With no supplied style, choose a restrained, legible direction appropriate to the footage and describe it briefly.

## Reusable examples

[FX authoring](fx-authoring.md) contains editable examples for footage, a card, grouped labels, and captions. [Motion](motion.md) shows keyframes and procedural animation. Adapt the composition and motion to the brief rather than copying the entire example look.
references/motion.md
# Motion

Use editable keyframes for specified poses, entrances, holds, and exits. Use
scripts for periodic motion or relationships that would otherwise require many
keys. Keep the structure editable: move a group for coordinated motion and its
children for independent details.

## Clocks and keyframes

`layerTime` is integer milliseconds from the owning layer's local start, not
project time. Ancestor group playback affects that clock. Keys outside the
visible interval may remain authored; the active window controls visibility.

This fades in group 3 from the FX authoring examples, holds it, then fades out:

```json
[
  {"type":"setFxPropertyKeyframes","compositionId":"main",
   "property":{"layerId":3,"propertyType":"opacity"},
   "keyframes":[
     {"id":"callout-opacity-in","layerTime":0,"value":{"type":"float","value":0},"easing":{"type":"linear"}},
     {"id":"callout-opacity-visible","layerTime":300,"value":{"type":"float","value":100},"easing":{"type":"cubicBezier","x1":0.2,"y1":0,"x2":0.2,"y2":1}},
     {"id":"callout-opacity-hold","layerTime":2500,"value":{"type":"float","value":100},"easing":{"type":"linear"}},
     {"id":"callout-opacity-out","layerTime":2900,"value":{"type":"float","value":0},"easing":{"type":"linear"}}
   ]}
]
```

Easing belongs to the destination key. Before/after the track the first/last
value is held. Keep key times distinct and IDs stable and composition-unique.
`setFxPropertyKeyframes` upserts by ID: omitted keys remain. Use
`removeFxPropertyKeyframes` to remove specific existing keys. Inspect the track
before revising it; do not accidentally append another animation over old keys.
Use `setFxPositionKeyframes` for paired spatial position paths with tangents.
Let the action API maintain the stored animation representation.

## Procedural motion

Check the installed schema before constructing a `setFxPropertyAnimator` action.
A `jsScript` body uses `layerTimeJsCode`, explicitly returns the target's value,
and reads `input.time.seconds` or `input.time.milliseconds`. For a scalar rotation
in degrees, a bounded oscillation body is:

```js
return Math.sin(input.time.seconds * Math.PI * 2) * 3;
```

Animation must evaluate correctly at arbitrary timestamps: do not accumulate
state frame by frame or use wall-clock time. Declare property dependencies when
following another property; dependency values are canonically ordered, not
necessarily in authored array order. Preserve existing relationships when editing.

## Design and review

Choose motion to support the requested emphasis. Keep text still long enough to
read; avoid moving every part independently. Establish the final layout first,
then add motion. Use easing for a deliberate arrival and linear interpolation
for constant-rate movement. Check overshoot against the available space.

Sample the entrance, settled hold, and exit, including just before and after a
transition. For the example, use `--timestamps-ms 0,150,300,1400,2500,2700,2900`.
These are project timestamps: offset them when the layer or its group starts
later. The exact end of a half-open active window is already invisible.

Review both the whole project and `--fx-solo main:3`. A content-bounds crop
helps inspect details but cannot prove correct canvas placement. Filmstrips
show sampled poses; inspect playback of the exported video for pacing, flicker,
and audio that sparse stills can miss.
references/native-authoring.md
# Authoring native Tesseract projects

## Local workflow

Use [local operation](local-operation.md) to create or inspect a `.tsrct`
document. `project checkout` writes editable JSON; `project commit` validates
and saves it. Supported composition-local action batches use `project apply`.
Keep the original and stable IDs. A rejected commit or action batch does not
publish partial changes. Never edit the archive manually or use a protobuf
converter for this document format.

Read [FX authoring](fx-authoring.md) before changing project structure, and
[motion](motion.md) before changing animation. These contain working examples;
use `project schema --document` and `project schema` for the installed fields.
The action schema includes only supported standalone operations and their field restrictions.
Use checkout/commit for canvas, duration, and media-layer structure changes.

## Essential invariants

1. `layers[0]` is topmost. Preserve sibling order when inserting footage behind graphics.
2. IDs are identities, not array positions. Layer IDs are unique throughout the composition's nested tree. Effect, mask, style, and text-animator IDs used by dynamics must remain stable.
3. New visual media uses `Video` or `Image`, not the `Media` tag. Import images with `project import-asset --kind image` before placing them; see [media import](media-import.md).
4. FX transform position is in parent pixels; anchor is layer-local. Scale and transform opacity are percentages, with `100` as identity. Color channels are `0..1`. Audio gain is linear. Check each field's schema rather than applying one universal percentage conversion.
5. The document owns one composition beginning at zero. Document `duration` is seconds; layer `activeRange` and `sourceRange` are milliseconds. Nested layer ranges are relative to the immediate parent. Moving a clip must not change its selected source moment.
6. `layerTimeJsCode` reads `input.time.seconds` from the owning layer's start. It must explicitly return a finite scalar/vector/color matching the target. Use stable IDs and declare dependencies. There is no DOM, CSS, or browser timeline.
7. Group related layers with supported grouping actions. Group effects see precomposed children; a normal layer effect sees only that layer. An Adjustment affects siblings below it. Inspect the schema before constructing a nontrivial feature.
8. Never reconstruct an unfamiliar document from a minimal template to make one correction: that drops existing layers, animation, and resources. Preserve unknown fields and re-checkout after actions to avoid committing stale JSON.

## Footage editing

Import each distinct local video once with `project import-video`. Select cuts
using Video layers: a clip from source 2.2–5.7s placed at edit 0–3.5s has
`sourceRange: {start: 2200, duration: 3500}` and
`activeRange: {start: 0, duration: 3500}` at composition root. Set
`sourceIntrinsicDuration` to the actual full source duration returned by import.
Set `volume: 1.0` on new Video layers to enable embedded source audio; an
omitted/null volume disables it. Preserve existing gain on revisions unless
deliberately changing the mix. Do not add a second
audible copy of the same sound. Use the Video layer's transform for a footage
push-in, or a Group transform when footage and graphics should move together.

For speed changes, maintain both source and edit duration deliberately. Inspect
the installed playback schema and render the result rather than inferring speed
behavior from different duration values. See [audio and timing](audio-and-timing.md).

## Assets and fonts

Resources are packaged in the `.tsrct` document. A filename or invented asset
ID in JSON does not import bytes. Use the ID returned by `project import-video`.
Choose and obtain fonts using the [text workflow](fx-authoring.md#text-with-an-available-font).
Import a local TTF/OTF/TTC with `project import-font --project project.tsrct
--file /path/to/font.ttf` before creating text. Use the returned `fontFamily`
and `fontStyle`. The command embeds
font bytes in the document. Use `project import-asset` for audio and images.
If rendering returns `{"error":"missing_fonts","fonts":[...]}`, import the
required files and retry. Do not silently replace the requested font.

Preview, filmstrip, and export use the same packaged font bytes. Check rendered
glyphs and spacing; a fallback font is not proof the intended face was used.
references/review-and-delivery.md
# Review and delivery

## Review a render, not an intention

There are three different checks:

1. **Structural:** commit/apply accepts the document; assets resolve; output is produced.
2. **Technical:** source frames and intervals are correct; alpha and audio are present as intended; duration, resolution, and codec match the brief; fonts render correctly.
3. **Creative:** the footage tells the intended story; graphics earn their space; pacing and sound work; the result matches the supplied reference and brand.

Passing the first two does not justify calling the result excellent. Avoid inflated ratings or unearned claims of professional quality. Be specific about what was reviewed.

## During editing

Use [filmstrip review](filmstrip-review.md) for visual revisions,
[waveform editing](waveform-editing.md) for music timing and dialogue cuts, and
[ad hooks](ad-hooks.md) when designing an opening. Apply the relevant procedure
to the changed beat and its neighbors before final review.

## What to inspect

- Opening: does the viewer see a meaningful subject/action, or just a generic title?
- Footage: correct source moments, intentional framing, complete actions, no accidental freezes/black gaps, continuity across cuts. For talking heads, check [eyeline continuity](editorial-decisions.md#talking-head-eyeline-continuity) in the final crop: stable eye height/screen position across cuts, natural head movement, and deliberate rather than accidental framing changes.
- Graphics: real hierarchy, sufficient reading time, consistent type/palette, intentional motion, clear value beyond repeating speech.
- Overlays: no collisions with faces, hands, product details, captions, or safe-area boundaries throughout the interval.
- Full scenes: spatial relationships and animation explain the idea; avoid a sequence of static slides with transition effects.
- Audio: speech clear over the bed, meaningful source sound retained, motivated cue density, no double playback or clipped transients, smooth ducks and intended dropouts, lip sync correct. Check mono and low-volume playback; measure the final encoded integrated loudness and true peak against the chosen spec.
- Ending: enough time to understand the final thought/CTA, no accidental audio cut, correct final frame.
- Editability: text and vector layers remain editable; source IDs and timing remain traceable; versioned files can be reopened.

Sample the entrance, readable hold, exit, and frames around every important cut. Then watch at speed with audio. A filmstrip helps coverage but cannot establish the feel of a cut or mix.

## Focused technical checks

Use local `ffprobe` to inspect the final stream and duration. Check that audio exists when it is expected; an AAC stream alone doesn't prove the waveform is non-silent. If needed inspect the mix with `volumedetect`/`ebur128` and listen.

For sound passes, use `python3 <skill-root>/scripts/tesseract_sound.py measure /absolute/path/final.mp4 --report /absolute/path/edit/.tesseract-work/checks/loudness-v1.json`. Use `mobile-preview` to make a separate mono, band-limited listening copy; never replace the deliverable with that diagnostic. Read [sound design](sound-design.md) for working targets and listening criteria.

After speech cleanup, compare the same original/processed phrases at matched loudness, then listen in the full mix. Check consonants, voice identity, breaths, room tails, processing artifacts, and lip sync; use [speech cleanup](speech-cleanup.md) for the procedure.

For transparent overlays, export ProRes 4444 MOV as described in [local operation](local-operation.md).
Extract a representative alpha frame with FFmpeg `-vf alphaextract` and inspect it;
verify both uncovered transparent pixels and the intended opaque content. Composite
over light and dark backgrounds to check edges and shadows. MP4 does not carry alpha.

For fonts, render representative text with the imported files and check it against
[the chosen typography](fx-authoring.md#text-with-an-available-font). Check both
visual fit and glyph/layout correctness; successful export alone proves neither.

## Handoff

Use the same project root established in [project lifecycle](local-operation.md#project-lifecycle).
Check that the current document, export, and previews correspond and packaged
resources resolve. Review the supporting artifacts: make useful previews visible,
keep retained authoring files within the working subfolder, and tidy disposable
scraps. Do not turn handoff into a second workspace containing selected copies.

Show the playable video and link the editable document, useful previews, and
project folder. Briefly describe the result or changes and any actual limitations.
Surface actionable review findings rather than listing internal files or raw
measurements. Do not describe generated clips as real footage or mocked workflows
as working integrations.
references/sound-design.md
# Sound design for social video and ads

Sound helps establish rhythm, direct attention, and make the edit feel intentional. Design it with the pictures, not as decoration added after every animation. Keep speech understandable on small speakers at low volume.

## Permission and source selection

When making a new edit or when asked to improve its sound, add appropriate local sounds without asking for each cue. Use the user's supplied audio, meaningful production sound from the footage, an available local sound library cleared for the project, or this plugin's procedural accents. Preserve an explicitly chosen track and obey requests for silence or a clean dialogue-only edit. A narrow visual revision does not authorize restyling the whole mix.

The sound helper offers eight procedural accents. List them with `python3 <skill-root>/scripts/tesseract_sound.py list` and use `make` to generate a WAV with the desired duration, seed, or level locally. These are synthesized design accents, not recordings of the actual product. Do not imply that a synthetic click is factual evidence of how something sounds. For documentary/product claims, prefer authentic source sound.

Import selected supplied or generated files with `project import-asset --kind audio`, then use separate native Audio layers and keep source/gain/timing traceable. Follow [audio and timing](audio-and-timing.md) for placement and mixing. Do not bake all sound into a replacement video or change the source footage. The local-only plugin does not contact a music, voice, or sound-generation service. If a bespoke soundtrack is essential and no suitable local track exists, identify that missing input; do not invent narration or silently replace the user's music.

For an ad opening, follow [ad hooks](ad-hooks.md) to map the chosen visual event to the music and a motivated sound accent. Align perceptual arrivals/transients, preserve speech, and use the existing music hit alone when it already provides the impact.

## Five roles to consider

These are five possible roles, not five compulsory tracks. Use only the layers that improve the piece; absence and silence are useful decisions.

| Role | What it contributes | How to mix and place it |
|---|---|---|
| Primary dialogue / voiceover | Meaning, character, narrative clarity | First priority. Keep the main voice centered; use clip gain and, when needed, compression for consistent intelligibility. Preserve breaths and natural emphasis. Heavy compression is a choice for uneven material, not an automatic preset. |
| Background music | Emotional direction, pacing, continuity | Establish a useful bed level, then duck it around speech. A gentle presence-region dip around 1–3 kHz can help if the music masks consonants; listen rather than hollowing out the track. Use lifts or dropouts to support a reveal or punchline. |
| Micro-SFX / UI accents | Points attention at a particular event | Pops, taps, clicks, paper sounds, and pings should match meaningful motion or interaction. Short upper-mid/high-frequency detail often reads on a phone; avoid piercing transients or repeated notification-like pings. |
| Structural transitions | Marks a change in scene, energy, or argument | Choose a whoosh, riser, downlifter, impact, or deliberate pause when the transition earns emphasis. Align the arrival/peak with the reveal, not automatically the sound file's beginning. Low impacts need audible harmonics so their meaning survives small speakers. |
| Diegetic foley / ambience | Gives footage physical context and continuity | Prefer source product clicks, handling, room tone, and environmental texture. For background texture, roughly 18–24 dB below dialogue is a starting relationship, not an absolute fader setting. Featured product sounds can come forward in speech gaps. |

For a fast ad, reassess attention every 2–4 seconds; that does not mean inserting a whoosh every 2–4 seconds. Avoid stacking a pop, hit, whoosh, and bass drop on the same trivial text entry. One deliberate cue often reads more clearly.

For speech that needs repair or tonal/level improvement, use [speech cleanup](speech-cleanup.md). Establish a natural, intelligible voice before building the mix around it; avoid amplifying room sound with excessive compression.

## Stack the mix around the voice

1. Establish the timing anchor. For talking-led work, tighten and verify the dialogue/source edit before mixing around it. When music was supplied or generated first, map its beats, phrases, and major changes before planning picture timing. Follow [waveform editing](waveform-editing.md) in either case; listen to confirm the waveform candidates. Speech remains intelligible and complete even in a music-led edit.
2. Place the music to establish the arc. Start a music duck around 3–6 dB below its unducked level, adjusting by ear to the actual tracks. Begin the reduction just before speech; use a short smooth attack and a slower release so the bed does not pump between syllables. As a starting point, try about 30–80 ms down and 150–350 ms back up.
3. Add micro-SFX only to actions the viewer should notice. Shape their tails and reduce them during important words.
4. Add structural cues to the few major changes. Check the sequence with those cues muted; they should improve the cut rather than conceal a weak one.
5. Restore subtle ambience/foley where it helps the footage feel connected. Retain useful production sound rather than burying it under added effects.

The relative duck is a multiplier: −3 dB ≈ 0.708 and −6 dB ≈ 0.501. Apply it to the selected bed gain. A fader value of 0.2 does not mean the bed is a fixed number of dB below the voice; the input waveforms determine that relationship.

## A brief drop in density

An 80–200 ms pause in music and background texture before a reveal can create contrast. Treat it as an option to audition, not a universal retention technique. Preserve any word or authentic sound carrying the meaning. If actual silence is intended, account for all audible layers and lingering tails; muting music alone may not create silence. Short 5–20 ms ramps can avoid clicks while still feeling abrupt. Let the reveal's cue/music return make the payoff legible.

## Frequency and mobile translation

- Protect the speech's presence, broadly in the mids. Start with level and arrangement before EQ; do not remove the entire 500 Hz–4 kHz region from every background track.
- Reserve deep low end for the music's bass/kick or a selected impact, and avoid overlapping several bass-heavy effects. Do not depend on sub-bass for information a phone listener must hear.
- Micro sounds can use upper-mid/high detail, often around 2–8 kHz, but brightness is not permission to boost everything above 5 kHz. Check sibilance and fatigue.
- Center the narrative voice and keep essential cues audible in mono. Check headphones, low playback volume, and a phone when available. The helper's mono/band-limited preview is a diagnostic aid, not an exact device simulation.

## Loudness and delivery

Use approximately −15 LUFS integrated, with an acceptable starting range of −16 to −14 LUFS, when there is no explicit delivery spec. This is a working social-video target, not a claim that every feed has the same normalization policy. Preserve intended contrast rather than forcing every short cue or quiet passage to that loudness.

Keep the **final encoded** true peak at or below **−1 dBTP**. True peak is measured in dBTP, not the sample-peak dBFS reading. Leaving a little extra margin before AAC encoding, such as −1.5 dBTP, can help, but remeasure the actual deliverable. Follow a supplied platform/client spec over this default.

Listen to the completed muxed video. Check speech clarity, cue density, abrupt edges, unintended silence, bass clutter, and the head/tail. Measure the final audio and keep the report with the project. A good meter reading is not proof of a good mix.

For exact local operations, native envelopes, and the distinction between timed ducking and signal-driven compression, read [audio and timing](audio-and-timing.md).
references/speech-cleanup.md
# Clean up talking audio

Read this when dialogue needs clearer tone, steadier levels, or less room noise/reverb. Apply only the processing the recording needs. Keep the speaker's identity, natural breathing, and complete words. A clean source may need only gain adjustment.

## Diagnose before choosing a treatment

Audition representative loud and quiet phrases, sibilant words, pauses, and phrase endings where reflections are audible. Check the source waveform and, when useful, a spectrum/spectrogram. A waveform helps with timing and levels; it cannot distinguish all noise, reflections, or intelligibility problems.

Separate **room tone** (background ambience/noise), **reverb** (the voice's decaying reflections), and **echo** (audible repeats). Also check for clipping, mic handling, plosives, hum, and changing mic distance. Lowering gain does not repair clipping already in the recording.

Probe and listen to available audio streams/channels first. A clean lavalier or boom track is preferable to repairing a distant camera mic. Do not blindly sum different mics: delayed copies can create comb filtering and more apparent room sound. Center a selected mono voice deliberately while preserving meaningful stereo production sound elsewhere.

## Choose a light treatment

These are starting points to audition, not a fixed chain or a requirement to use every processor.

| Problem | Useful first treatment | What to listen for |
|---|---|---|
| Uneven speaking levels | Ride phrase/clip gain before relying on heavy compression; preserve intentional emphasis. | The voice should stay present without making breaths and room tails unnaturally loud. Peak normalization alone does not level speech. |
| Rumble or handling noise | Try a gentle high-pass around 60–90 Hz, lowering the cutoff for a voice with useful bass. Address isolated plosives locally when possible. | Stop before the voice loses body. A high-pass is not a reverb remover. |
| Mud or boxiness | Find the actual resonance; audition a broad 1–3 dB reduction somewhere around 200–500 Hz. | Avoid automatically scooping all low mids or making the voice thin. Room coloration can occur outside this range. |
| Harshness or poor intelligibility | Reduce the offending band modestly before adding brightness. Keep room for speech by reducing competing music. | A presence boost around 2–5 kHz can also amplify harshness, hiss, and reflections; clearer is not always brighter. |
| Steady hiss, fan noise, or hum | Use restrained noise reduction; try roughly 3–6 dB first. A narrow hum notch is appropriate only at a confirmed hum frequency/harmonic. | Listen for watery/bubbly artifacts and lost consonants. Learn a noise profile only from an actual speech-free segment; a changing background is not stationary noise. |
| Excess dynamics after gain rides | Try a soft-knee compressor around 2:1–3:1, attack 10–30 ms, release 100–250 ms. Set threshold from this recording; a few dB of reduction on louder phrases is a useful first audition. | Too-fast attack dulls consonants; a fast release or excessive makeup gain can bring up room sound and cause pumping. Adjust by phrase behavior, not a preset label. |
| Sharp “s” and “sh” sounds | Apply selective de-essing in the actual sibilant region, often roughly 4–10 kHz. Recheck after EQ/compression. | Preserve articulation. Excess de-essing creates a lisp; a permanent large treble cut dulls the whole voice. |
| Audible noise between phrases | Prefer natural room tone, gentle gain rides, or mild expansion when needed. | Hard gates can chop breaths and word tails and make the room switch on/off. Silence detection is not word-boundary detection. |

A useful order to audition is source/channel selection → repair or light noise reduction → corrective EQ → leveling/compression → selective de-essing → mix gain. Change the order or omit stages according to the recording. Avoid stacking several enhancement passes that each remove speech detail. Use [waveform editing](waveform-editing.md) when shortening pauses; processing does not make an unsafe trim safe.

## Room reverb and echo

EQ may reduce room boxiness, but cannot separate all reflections from the direct voice. Gating reduces tails in gaps and does not remove reverb underneath words. Heavy compression often makes the problem more noticeable. Retain some room character if removing it damages the voice.

This plugin has **no bundled dedicated dereverberation or speech-isolation model**. FFmpeg's `afftdn` is a noise reducer; `aecho` adds echo. Neither is a general speech dereverberator. If a suitable local enhancement tool/model is already available and permitted, audition it conservatively on a copy. Do not silently download a model or upload dialogue to an external service. Echo cancellation that requires a clean reference signal is not a general solution when that reference is absent.

Compare a moderate treatment before increasing strength. Check phrase endings, quiet consonants, natural breaths, and voice texture. Back off if it sounds metallic, watery, phasey, or syllables disappear. For a badly reverberant recording, report the remaining limitation honestly; improved intelligibility is a useful result even when the room cannot be fully removed. Missing words and a clean studio recording cannot be guaranteed from damaged source audio.

## Apply locally and preserve timing

Tesseract's native controls are source/active ranges, gain, and gain envelopes. EQ, compression, denoising, and de-essing use a **new local audio derivative**, not invented Tesseract DSP fields. Inspect installed filter options with `ffmpeg -h filter=<name>` before using them. In particular, `acompressor` threshold/makeup use linear amplitude and attack/release use milliseconds; FFmpeg `deesser`'s `f` parameter is normalized, not a frequency in Hz.

This optional example auditions gentle rumble/boxiness reduction and compression on an **already aligned dialogue-only WAV**. Choose its settings from the actual recording, or omit unnecessary filters:

```bash
ffmpeg -nostdin -n -i /absolute/path/edit/dialogue-aligned.wav -map 0:a:0 -vn -af "highpass=f=70:p=2,equalizer=f=300:t=q:w=1:g=-2,acompressor=threshold=0.125:ratio=2:attack=15:release=150:makeup=1" -ar 48000 -c:a pcm_s24le /absolute/path/edit/dialogue-clean-v1.wav
```

`threshold=0.125` is approximately −18 dBFS and may be inappropriate for another input level. This example performs no de-reverb, loudness normalization, or true-peak limiting. For steady noise, the locally available `afftdn` can be auditioned separately with modest `nr` and a noise floor derived from the recording; do not call its defaults a universal cleanup preset. Its noise-only output mode can help reveal whether removed material includes words.

Keep the original file and record the selected stream/channel, processing recipe, source range, and mapping to project time. Favor processing continuous dialogue takes before trimming into many short clips; independent per-clip processing can create changing noise floors and processor transients. If processing an excerpt, include handles and record its offset.

Verify duration, head/tail, channel layout, processing latency, and lip sync before replacing audio in the project. A decoded WAV may start at zero even when its source video's audio starts later. Preserve that offset explicitly. Never time-stretch or trim the source to make a derivative appear aligned. Import the processed derivative with `project import-asset --kind audio` using a new asset ID. Then replace only the intended dialogue or mute its original playback when using an independent Audio layer, avoiding doubled speech. Keep music and SFX separate.

## Decide whether it is actually better

Compare original and processed versions over the **same phrases at matched perceived loudness**. Use integrated/short-term measurements to establish a comparison level, then listen; peak matching alone can bias the comparison. Audition quiet phrases, loud words, “s” sounds, and room tails with music muted, then in the full mix. Change one major treatment at a time when diagnosing artifacts.

Keep the milder version if it preserves more speech detail. Check headphones, mono, and low-volume playback; verify waveform joins and lip sync after processing. Follow [audio and timing](audio-and-timing.md) for final encoded loudness/true-peak checks. Do not normalize each voice clip independently to the complete-program target.

Record the actual benefit and any remaining room sound or artifacts. If audio cannot be auditioned in the host, processing and meter results remain unverified for listening quality; do not describe them as successful cleanup.
references/waveform-editing.md
# Edit with the audio waveform

Use the waveform as an editing view, alongside listening and the filmstrip. When music is supplied or generated first, inspect it **before deciding picture and motion timing**. When trimming talking footage, inspect its source audio **before cutting** and check the rendered joins afterward. Generating the image is only preparation: open it with the host's image-view tool and record what it shows.

## Generate an overview, then zoom

The local sound helper accepts audio files or the audio inside a video. It writes a labeled PNG and a JSON sidecar containing 10 ms per-channel peak/RMS measurements and candidate markers. It does not alter the source, transcribe speech, or edit a project.

```bash
# Inspect the selected music before planning the visual beats.
python3 "$SKILL/scripts/tesseract_sound.py" waveform /absolute/path/music.wav --mode music --output /absolute/path/edit/.tesseract-work/checks/music-overview.png

# Inspect a range of talking footage; times remain relative to the SOURCE file.
python3 "$SKILL/scripts/tesseract_sound.py" waveform /absolute/path/talking.mov --mode speech --start-ms 12000 --duration-ms 8000 --output /absolute/path/edit/.tesseract-work/checks/speech-source.png

# Zoom around proposed source cuts, keeping the original timestamps on the axis.
python3 "$SKILL/scripts/tesseract_sound.py" waveform /absolute/path/talking.mov --mode speech --start-ms 14500 --duration-ms 1800 --markers-ms 15080,15720 --output /absolute/path/edit/.tesseract-work/checks/speech-cut-check.png

# Inspect the RENDERED edit around the new join, on the EDIT file's clock.
python3 "$SKILL/scripts/tesseract_sound.py" waveform /absolute/path/edit/dialogue-v2.wav --mode speech --clock edit --start-ms 2000 --duration-ms 2000 --markers-ms 3100 --output /absolute/path/edit/.tesseract-work/checks/join-v2.png
```

Choose ranges and markers from the actual material; the times above are syntax examples. Default is the whole input, up to ten minutes per view. Split longer material into labeled ranges; even shorter overviews need zooms for word boundaries. Use `--audio-stream N` to select a zero-based audio-stream ordinal after probing files with multiple audio tracks. Prefer an isolated dialogue track over a camera scratch mix with background music.

The display separates channels, uses a square-root amplitude scale to expose quieter details, and does not normalize gain. Gold lines are transient/energy-rise candidates, pink lines are approximate energy changes, green bars are low-energy intervals, and white lines are the requested review markers. The plot caps each candidate category at 100 markers for legibility; all candidates remain in JSON. Threshold comparisons use the original decoded sample levels, not display height. Numeric timestamps in JSON are authoritative; use the image to see their context.

A `--clock edit` label means the input is an exported edit or stem. It does **not** convert source time to project time. A short preview file begins at zero on its own clock; add its project-range offset explicitly. The helper preserves delayed audio relative to the container start and fills timestamp gaps for inspection. A requested window can end at audio EOF; check the reported actual window. No audio stream is a missing input, not a silent waveform.

## Music first: give the edit a musical structure

1. Open the overview and listen to the selected track. Identify the first usable downbeat, recurring accents, phrase lengths, builds, breaks/dropouts, drops, arrangement changes, and the ending. The waveform helps locate them; a loud transient might be a snare or sound effect, and a new section can begin without getting louder.
2. Use the helper's accent and energy-change candidates as places to investigate. Zoom around important arrivals and audition them. Confirm a beat grid over several bars if useful; check for half/double-time errors, swing, pickups, and tempo changes. Do not assume every peak is a beat or infer a definitive BPM/section map from these heuristics. A dense mastered track may need listening to establish most of the map.
3. Save a small `audio-map.json` or table with music-source time, edit time, confirmed event, and intended visual action. Plan the opening, scene changes, product reveal, motion arrival, and ending against this map. Choose a few meaningful accents; cuts on every beat quickly become mechanical. Motion can start before a beat so its arrival lands on it.
4. Keep the chosen music and its recognizable phrases intact unless the brief calls for a music edit. When shortening it, cut at compatible phrase boundaries, audition the join, and preserve release/reverb tails. Do not globally time-stretch footage or speech to force alignment.
5. Recheck the picture against the music after timing changes. Pair the audio map with filmstrip samples, then play the assembled video with sound; a still image cannot prove synchronization or rhythm.

For music played at normal speed, `edit_time = layer_edit_start + music_source_time - selected_music_source_start`. For a constant source playback rate `r`, divide that source-time difference by `r`. Variable speed requires the actual time mapping; do not use the simple formula across a speed ramp. Preserve separate source, project, and layer clocks when authoring native timing.

## Talking footage: tighten pauses without damaging words

1. Audition the source and read any supplied transcript. Identify complete thoughts, meaningful emphasis, and pauses that can be shortened. Make a speech waveform view before changing source ranges. Never fabricate a transcript or assume quiet pixels mean no speech.
2. Inspect candidate pauses. By default, a green interval means **all channels** stay below a −45 dBFS sample-peak threshold for at least 200 ms, measured in 10 ms bins. Adjust `--quiet-db` and `--min-quiet-ms` for the recording; these are candidate settings, not universal speech thresholds. Quiet consonants, trailing syllables, breaths, and room tone can sit below a threshold. Constant noise or music can hide real pauses.
3. Zoom to roughly 0.5–2 seconds around each intended boundary and listen to the surrounding words. Choose trims in verified gaps, retaining enough lead-in and tail for consonants and natural breathing. Approximately 50–120 ms of handles is a useful first audition, not guaranteed protection. Keep longer pauses that carry meaning. Shorten an awkward pause rather than automatically deleting all silence.
4. Update video and its linked source audio together. Preserve source in/out points and edit positions. At video-frame boundaries, favor retaining a little more material over clipping a word. Use a deliberate J/L cut or room-tone bridge when it improves continuity, with source handles intact. Tiny fades may prevent clicks; fades and crossfades cannot restore missing phonemes and must not blur adjacent words.
5. Render the changed dialogue and inspect its waveform around every changed join, marked on the edit clock. Listen to the previous word, the join, and the next word at normal speed, with music muted first; inspect video frames for jump cuts and lip sync. Then listen in the full mix. Check word completeness, breathing, pacing, clicks, doubled syllables, and any new silence. Correct and re-render the affected range before continuing.

Use a temporary `.tsrct` copy with music/SFX gains and gain animators muted, export a dialogue-only MP4 with `tsrct export`, and pass that MP4 directly to the waveform helper. Preserve the working mix and original source; the CLI has no audio-only export command. For very short preview exports, retain the offset from preview time to full project time in the review notes.

## Record the actual checks

Record the revision, source file/stream, inspected ranges, confirmed music events
or speech boundaries, source-to-edit mapping, and what changed after review.
Waveforms support visual and auditory judgment; they do not establish intact words or a good mix on their own. If the host cannot audition audio, state that speech cuts and musical timing remain unverified instead of claiming they passed.

All analysis is local with Python's standard library and local FFmpeg/ffprobe. PNG labeling uses FFmpeg's `drawtext` filter and a locally available diagnostic font; project typography still uses fonts packaged with `project import-font`. No new service, model, upload, or paid generation is involved.
scripts/tesseract_media.py
"""Local file and FFmpeg helpers used by the sound and waveform tools."""
import hashlib
import json
import os
from pathlib import Path
import shutil
import subprocess
import sys
import tempfile


def emit(value):
    print(json.dumps(value, indent=2, ensure_ascii=False))


def sha256(path):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        for chunk in iter(lambda: f.read(1024 * 1024), b""):
            h.update(chunk)
    return h.hexdigest()


def existing(value):
    p = Path(value).expanduser().resolve()
    if not p.is_file():
        raise ValueError(f"Local file does not exist: {p}")
    return p


def external(name):
    p = shutil.which(name)
    if p:
        return p
    for parent in ("/opt/homebrew/bin", "/usr/local/bin"):
        p = Path(parent) / name
        if p.is_file():
            return str(p)
    raise ValueError(f"{name} is needed for media inspection. It was not found locally.")


def run(command, timeout=600):
    result = subprocess.run([str(x) for x in command], capture_output=True, text=True, timeout=timeout)
    if result.returncode:
        raise RuntimeError(f"{Path(str(command[0])).name} failed ({result.returncode}):\n{result.stderr[-6000:]}\n{result.stdout[-2000:]}")
    if result.stderr:
        print(result.stderr[-4000:], file=sys.stderr)
    return result.stdout


def output_path(value):
    p = Path(value).expanduser().resolve()
    if p.exists():
        raise ValueError(f"Output already exists: {p}. Choose a new versioned filename.")
    p.parent.mkdir(parents=True, exist_ok=True)
    return p


def atomic_run(output, command):
    p = output_path(output)
    with tempfile.TemporaryDirectory(prefix=".tesseract-", dir=p.parent) as tmp:
        temporary = Path(tmp) / p.name
        run(command(temporary))
        if not temporary.is_file() or not temporary.stat().st_size:
            raise RuntimeError("Media command returned without producing a nonempty output.")
        # Do not overwrite a concurrently-created output.
        os.link(temporary, p)
    return p


def write_json(path, data):
    p = output_path(path)
    with p.open("x") as f:
        json.dump(data, f, indent=2, ensure_ascii=False)
        f.write("\n")
    return p
scripts/tesseract_sound.py
#!/usr/bin/env python3
"""Local sound accents and mix inspection; no account, download, or paid service."""
import argparse
from array import array
import json
import math
from pathlib import Path
import random
import re
import subprocess
import sys
import wave

import tesseract_media as local
import tesseract_waveform

RATE = 48000
PRESETS = {
    "soft-tap": (100, "micro", "Muted tactile accent for a button or product contact."),
    "dry-click": (65, "micro", "Short bright tick for a precise graphic event."),
    "soft-pop": (180, "micro", "Rounded accent for a small reveal; avoid every text entry."),
    "clear-ping": (480, "micro", "Light two-partial tone for a meaningful confirmation."),
    "air-whoosh": (450, "transition", "Filtered air sweep; align its peak to the movement."),
    "short-riser": (1100, "transition", "Brief building texture into a reveal, with a clean end."),
    "downlifter": (800, "transition", "Descending texture after a scene change."),
    "soft-impact": (400, "transition", "Low hit with audible upper harmonics for small speakers."),
}


def synthesize(name, duration_ms, peak_db=-12.0, seed=0):
    if name not in PRESETS or not 40 <= duration_ms <= 4000:
        raise ValueError("Use a listed preset and a duration from 40 to 4000 ms.")
    if not math.isfinite(peak_db) or not -60 <= peak_db <= -3:
        raise ValueError("Cue peak must be between -60 and -3 dBFS; leave mix headroom.")
    rng = random.Random(seed)
    count = round(RATE * duration_ms / 1000)
    samples, phase, smooth = [], 0.0, 0.0
    duration = count / RATE
    for i in range(count):
        t = i / RATE
        u = i / (count - 1)
        noise = rng.uniform(-1.0, 1.0)
        smooth += 0.12 * (noise - smooth)
        bright = noise - smooth
        if name == "soft-tap":
            value = (0.8 * smooth + 0.2 * math.sin(2 * math.pi * 420 * t)) * math.exp(-8 * u)
        elif name == "dry-click":
            value = bright * math.exp(-15 * u)
        elif name == "soft-pop":
            phase += 2 * math.pi * (210 + 850 * math.exp(-20 * u)) / RATE
            value = (math.sin(phase) + 0.13 * bright) * math.exp(-7 * u)
        elif name == "clear-ping":
            value = (math.sin(2 * math.pi * 1100 * t) + 0.25 * math.sin(2 * math.pi * 1650 * t)) * math.exp(-6 * u)
        elif name == "soft-impact":
            phase += 2 * math.pi * (65 + 170 * math.exp(-12 * u)) / RATE
            value = (math.sin(phase) + 0.3 * math.sin(2 * phase) + 0.25 * bright * math.exp(-35 * u)) * math.exp(-7 * u)
        else:
            amount = u if name == "short-riser" else 1-u if name == "downlifter" else math.sin(math.pi * u)
            phase += 2 * math.pi * (500 + 1900 * amount * amount) / RATE
            value = (0.7 * smooth + 0.12 * bright + 0.06 * math.sin(phase)) * amount ** 1.7
        # Brief ramps remove hard waveform discontinuities without dulling the cue.
        fade = min(1.0, t / 0.002, (duration - 1/RATE - t) / 0.008)
        samples.append(value * max(0.0, fade))
    peak = max(abs(x) for x in samples)
    scale = 10 ** (peak_db / 20) / peak
    return array("h", (round(x * scale * 32767) for x in samples))


def save_cue(name, path, duration_ms=None, peak_db=-12.0, seed=0):
    duration_ms = PRESETS[name][0] if duration_ms is None else duration_ms
    samples = synthesize(name, duration_ms, peak_db, seed)
    output = local.output_path(path)
    if output.suffix.lower() != ".wav":
        raise ValueError("Local procedural cues are PCM WAV files; use .wav.")
    if sys.byteorder != "little":
        samples.byteswap()
    with output.open("xb") as f:
        with wave.open(f, "wb") as w:
            w.setnchannels(1)
            w.setsampwidth(2)
            w.setframerate(RATE)
            w.writeframes(samples.tobytes())
    return {"path": str(output), "preset": name, "duration_ms": duration_ms,
            "sample_peak_dbfs": peak_db, "seed": seed,
            "origin": "Locally synthesized from oscillators and seeded noise; no recorded samples."}


def measure(path):
    src = local.existing(path)
    result = subprocess.run([local.external("ffmpeg"), "-hide_banner", "-nostdin", "-i", str(src),
        "-map", "0:a:0", "-af", "loudnorm=I=-15:TP=-1:LRA=7:print_format=json",
        "-f", "null", "-"], capture_output=True, text=True, timeout=600)
    if result.returncode:
        raise RuntimeError(result.stderr[-2500:])
    candidates = re.findall(r'\{\s*"input_i".*?\}', result.stderr, re.S)
    if not candidates:
        raise RuntimeError("FFmpeg did not return loudness measurements.")
    raw = json.loads(candidates[-1])
    def finite(key):
        v = float(raw[key])
        return v if math.isfinite(v) else None
    il, tp = finite("input_i"), finite("input_tp")
    return {"path": str(src), "audio_stream": "first audio stream", "integrated_lufs": il,
            "true_peak_dbtp": tp, "loudness_range_lu": finite("input_lra"),
            "working_target": {"integrated_lufs": [-16, -14], "true_peak_max_dbtp": -1},
            "within_working_target": il is not None and tp is not None and -16 <= il <= -14 and tp <= -1,
            "note": "Measurement only. Targets are a social-edit starting point, not a platform guarantee. Listen to the mix; silence/very short cues may not have meaningful integrated loudness."}


def main():
    parser = argparse.ArgumentParser(description=__doc__)
    sub = parser.add_subparsers(dest="command", required=True)
    sub.add_parser("list")
    p = sub.add_parser("make")
    p.add_argument("--preset", choices=PRESETS, required=True)
    p.add_argument("--output", required=True)
    p.add_argument("--duration-ms", type=int)
    p.add_argument("--peak-db", type=float, default=-12)
    p.add_argument("--seed", type=int, default=0)
    p = sub.add_parser("measure")
    p.add_argument("input")
    p.add_argument("--report")
    p = sub.add_parser("waveform", help="Labeled waveform PNG + JSON candidates for music or speech review")
    p.add_argument("input")
    p.add_argument("--output", required=True)
    p.add_argument("--report")
    p.add_argument("--mode", choices=("music", "speech"), required=True)
    p.add_argument("--start-ms", type=int, default=0)
    p.add_argument("--duration-ms", type=int)
    p.add_argument("--quiet-db", type=float, default=-45)
    p.add_argument("--min-quiet-ms", type=int, default=200)
    p.add_argument("--markers-ms", help="Comma-separated times on the input file clock; mark proposed cuts or accents")
    p.add_argument("--clock", choices=("source", "edit"), default="source")
    p.add_argument("--audio-stream", type=int, default=0, help="Zero-based audio stream ordinal")
    p = sub.add_parser("mobile-preview")
    p.add_argument("input")
    p.add_argument("--output", required=True)
    args = parser.parse_args()
    if args.command == "list":
        local.emit([{"preset": k, "duration_ms": v[0], "role": v[1], "use": v[2]} for k,v in PRESETS.items()])
    elif args.command == "make":
        local.emit(save_cue(args.preset, args.output, args.duration_ms, args.peak_db, args.seed))
    elif args.command == "measure":
        report = measure(args.input)
        if args.report:
            local.write_json(args.report, report)
        local.emit(report)
    elif args.command == "waveform":
        markers = [float(x) for x in args.markers_ms.split(",")] if args.markers_ms else []
        local.emit(tesseract_waveform.waveform(args.input, args.output, args.report, args.mode,
            args.start_ms, args.duration_ms, args.quiet_db, args.min_quiet_ms, markers,
            args.clock, args.audio_stream))
    else:
        src = local.existing(args.input)
        if Path(args.output).suffix.lower() != ".wav":
            raise ValueError("Use .wav for the diagnostic listening copy.")
        p = local.atomic_run(args.output, lambda tmp: [local.external("ffmpeg"), "-hide_banner", "-nostdin", "-n",
            "-i", src, "-map", "0:a:0", "-vn", "-af", "aformat=channel_layouts=mono,highpass=f=150,lowpass=f=7000",
            "-ar", "48000", "-c:a", "pcm_s16le", tmp])
        local.emit({"output": str(p), "purpose": "Mono, band-limited diagnostic copy. Not the deliverable or an exact phone simulation."})
    return 0


if __name__ == "__main__":
    try:
        raise SystemExit(main())
    except (ValueError, RuntimeError, OSError, subprocess.TimeoutExpired) as error:
        print(f"tesseract-sound: {error}", file=sys.stderr)
        raise SystemExit(1)
scripts/tesseract_waveform.py
#!/usr/bin/env python3
"""Timestamped local waveform views and review candidates; never automatic edits."""
from array import array
import json
import math
import os
from pathlib import Path
import statistics
import sys
import tempfile

import tesseract_media as local

RATE = 48000
BIN_MS = 10
MAX_WINDOW_MS = 600000


def db(value):
    return round(20 * math.log10(value), 2) if value > 0 else None


def envelope(raw, channels, start_ms):
    """Keep channels separate: anti-phase stereo must not look like silence."""
    points = []
    frames = 0
    with Path(raw).open('rb') as f:
        while True:
            block = f.read(RATE * BIN_MS // 1000 * channels * 4)
            if not block:
                break
            if len(block) % (channels * 4):
                raise ValueError('Incomplete decoded audio frame.')
            samples = array('f', block)
            if sys.byteorder != 'little':
                samples.byteswap()
            if not all(math.isfinite(s) for s in samples):
                raise ValueError('Audio contains non-finite samples.')
            count = len(samples) // channels
            peaks, rms = [], []
            for c in range(channels):
                values = samples[c::channels]
                peaks.append(max(abs(s) for s in values))
                rms.append(math.sqrt(sum(s*s for s in values) / count))
            points.append({'start_ms': round(start_ms + frames * 1000 / RATE, 3),
                           'end_ms': round(start_ms + (frames + count) * 1000 / RATE, 3),
                           'channel_peak_dbfs': [db(v) for v in peaks],
                           'channel_rms_dbfs': [db(v) for v in rms]})
            frames += count
    if not points:
        raise ValueError('No audio in the requested window.')
    return points, frames


def peak(point):
    return max((v for v in point['channel_peak_dbfs'] if v is not None), default=-160)


def rms(point):
    return max((v for v in point['channel_rms_dbfs'] if v is not None), default=-160)


def candidates(points, mode, quiet_db, min_quiet_ms):
    quiet, begin = [], None
    for i, p in enumerate(points):
        low = peak(p) < quiet_db
        if low and begin is None:
            begin = p['start_ms']
        if begin is not None and (not low or i == len(points)-1):
            end = p['end_ms'] if low else p['start_ms']
            if end - begin >= min_quiet_ms:
                quiet.append({'start_ms': begin, 'end_ms': end})
            begin = None
    result = {'quiet_intervals': quiet}
    if mode != 'music':
        return result
    # Abrupt short-window energy rises, not beat/downbeat or tempo estimation.
    levels = [rms(p) for p in points]
    accents = []
    for i in range(1, len(points)):
        previous = statistics.median(levels[max(0, i-10):i])
        rise = levels[i] - max(previous, -80)
        if levels[i] > -45 and rise >= 6 and levels[i] - levels[i-1] >= 3:
            if not accents or points[i]['start_ms'] - accents[-1]['time_ms'] >= 120:
                accents.append({'time_ms': points[i]['start_ms'], 'rise_db': round(rise, 2)})
    # Compare one-second neighborhoods; do not label these as semantic sections.
    changes = []
    for i in range(100, len(points)-100, 25):
        before = statistics.mean(max(v, -80) for v in levels[i-100:i])
        after = statistics.mean(max(v, -80) for v in levels[i:i+100])
        delta = after - before
        if abs(delta) >= 6:
            changes.append({'time_ms': points[i]['start_ms'], 'energy_change_db': round(delta, 2)})
    selected = []
    for candidate in sorted(changes, key=lambda v: -abs(v['energy_change_db'])):
        if all(abs(candidate['time_ms'] - v['time_ms']) >= 1500 for v in selected):
            selected.append(candidate)
    changes = sorted(selected, key=lambda v: v['time_ms'])
    result.update({'accent_candidates': accents, 'energy_change_candidates': changes})
    return result


def draw_filter(report):
    width, height, left, top = 1320, 360, 60, 86
    start, end = report['window_start_ms'], report['window_end_ms']
    def x(t):
        return left + min(width-1, max(0, round(width * (t-start) / (end-start))))
    diagnostic_font = Path('/System/Library/Fonts/Supplemental/Arial.ttf')
    font_option = f"fontfile='{diagnostic_font}':" if diagnostic_font.is_file() else ''
    def label(text, px, py, size=18, color='0xe5e7eb'):
        # Only internally generated labels/numbers are passed into filter syntax.
        return f"drawtext={font_option}text='{text}':expansion=none:fontsize={size}:fontcolor={color}:x={px}:y={py}"
    filters = [f'showwavespic=s={width}x{height}:split_channels=1:colors=0x7dd3fc|0xc4b5fd|0x86efac|0xfda4af|0xfde68a|0x67e8f9|0xd8b4fe|0xfdba74:scale=sqrt:filter=peak',
               'format=rgb24', f'pad=1440:560:{left}:{top}:color=0x111827']
    for i in range(7):
        px = left + round((width-1)*i/6)
        t = start + (end-start)*i/6
        filters.append(f'drawbox=x={px}:y={top}:w=1:h={height}:[email protected]:t=fill')
        filters.append(label(f'{t/1000:.3f}s', f'{px}-text_w/2', top+height+12, 17))
    marks = report['candidates']
    # Overview markers are capped for readability; JSON keeps every candidate.
    for q in marks['quiet_intervals'][:100]:
        filters.append(f"drawbox=x={x(q['start_ms'])}:y={top+height-8}:w={max(1,x(q['end_ms'])-x(q['start_ms']))}:h=8:color=0x86efac:t=fill")
    for q in marks.get('accent_candidates', [])[:100]:
        filters.append(f"drawbox=x={x(q['time_ms'])}:y={top}:w=1:h={height}:[email protected]:t=fill")
    for q in marks.get('energy_change_candidates', [])[:100]:
        filters.append(f"drawbox=x={x(q['time_ms'])}:y={top}:w=2:h={height}:[email protected]:t=fill")
    for t in report['review_markers_ms']:
        filters.append(f'drawbox=x={x(t)}:y={top}:w=2:h={height}:color=white:t=fill')
    for c in range(report['channels']):
        filters.append(label(f'CH {c+1}', left+8, top+c*height//report['channels']+6, 13))
    filters.append(label(f"{report['mode'].upper()} WAVEFORM - {report['clock'].upper()} TIME", 60, 22, 24))
    filters.append(label('Seconds from media start / separate channels / square-root amplitude / no gain normalization', 60, 57, 16, '0x9ca3af'))
    if report['mode'] == 'music':
        legend = 'Gold = accent candidates   Pink = energy shifts   Green = quiet intervals   White = review markers'
    else:
        legend = 'Green = quiet candidates across ALL channels   White = review markers   Quiet does not mean safe to cut'
    filters.append(label(legend, 60, 492, 17))
    filters.append(label('Open and inspect this image. Listen before confirming a beat or edit. Zoom around speech cuts.', 60, 525, 16, '0x9ca3af'))
    return ','.join(filters)


def waveform(path, output, report_path=None, mode='music', start_ms=0, duration_ms=None,
             quiet_db=-45, min_quiet_ms=200, markers_ms=(), clock='source', audio_stream=0):
    if mode not in ('music', 'speech') or clock not in ('source', 'edit'):
        raise ValueError('Use music/speech mode and source/edit clock.')
    if start_ms < 0 or not math.isfinite(quiet_db) or not -90 <= quiet_db <= -15:
        raise ValueError('Start must be nonnegative; quiet threshold must be -90 to -15 dBFS.')
    if min_quiet_ms < BIN_MS or min_quiet_ms > MAX_WINDOW_MS:
        raise ValueError('Quiet duration must be 10 to 600000 ms.')
    src = local.existing(path)
    out = local.output_path(output)
    if out.suffix.lower() != '.png':
        raise ValueError('Use .png for the waveform view.')
    dest = local.output_path(report_path or out.with_suffix('.json'))
    if dest == out or dest.suffix.lower() != '.json':
        raise ValueError('Use a distinct .json report path.')
    meta = json.loads(local.run([local.external('ffprobe'), '-v', 'error', '-show_streams', '-show_format', '-of', 'json', src], 30))
    streams = [s for s in meta.get('streams', []) if s.get('codec_type') == 'audio']
    if audio_stream < 0 or audio_stream >= len(streams):
        raise ValueError('Requested audio stream does not exist; probe the file first.')
    stream = streams[audio_stream]
    channels = int(stream['channels'])
    if not 1 <= channels <= 8:
        raise ValueError('Waveform views support 1 to 8 channels per audio stream.')
    fmt = meta.get('format', {})
    origin = float(fmt.get('start_time') or 0)
    media_duration = float(fmt['duration']) * 1000 if fmt.get('duration') else None
    if duration_ms is None:
        if media_duration is None:
            raise ValueError('Duration unknown; provide --duration-ms.')
        duration_ms = math.ceil(media_duration) - start_ms
    if not 0 < duration_ms <= MAX_WINDOW_MS:
        raise ValueError('Choose a window of 1 to 600000 ms; split long media into labeled ranges.')
    if media_duration is not None and start_ms >= media_duration:
        raise ValueError('Window starts after the end of the media.')
    markers = sorted(set(markers_ms))
    if len(markers) > 48 or any(not math.isfinite(t) or not start_ms <= t < start_ms+duration_ms for t in markers):
        raise ValueError('Use up to 48 review markers within the requested window.')
    with tempfile.TemporaryDirectory(prefix='.tesseract-waveform-', dir=out.parent) as tmp:
        raw = Path(tmp)/'segment.f32'
        # Preserve timestamps relative to the container origin, including delayed audio.
        # Resampling fills timestamp gaps; no mono summing or level normalization occurs.
        af = (f'asetpts=PTS-({origin:.9f})/TB,aresample={RATE}:async=1:first_pts=0,'
              f'atrim=start_sample={start_ms*RATE//1000}:end_sample={(start_ms+duration_ms)*RATE//1000},asetpts=PTS-STARTPTS')
        local.run([local.external('ffmpeg'), '-v', 'error', '-nostdin', '-n', '-copyts', '-i', src,
                   '-map', f'0:a:{audio_stream}', '-vn', '-af', af, '-t', f'{duration_ms/1000:.6f}',
                   '-c:a', 'pcm_f32le', '-f', 'f32le', raw])
        points, frames = envelope(raw, channels, start_ms)
        end = round(start_ms+frames*1000/RATE, 3)
        if any(t >= end for t in markers):
            raise ValueError('A review marker is beyond the decoded audio window.')
        report = {'input': str(src), 'input_sha256': local.sha256(src), 'mode': mode, 'clock': clock,
                  'audio_stream_ordinal': audio_stream, 'container_origin_seconds': origin,
                  'channels': channels, 'analysis_sample_rate': RATE, 'envelope_bin_ms': BIN_MS,
                  'window_start_ms': start_ms, 'window_end_ms': end,
                  'requested_duration_ms': duration_ms, 'decoded_duration_ms': round(frames*1000/RATE, 3),
                  'quiet_threshold_dbfs': quiet_db, 'min_quiet_ms': min_quiet_ms,
                  'review_markers_ms': markers, 'image': str(out), 'report': str(dest),
                  'candidates': candidates(points, mode, quiet_db, min_quiet_ms), 'envelope': points,
                  'interpretation': 'All timestamps are milliseconds from the input media start on the labeled clock. Quiet intervals require every channel below the peak threshold. Accent and energy candidates are heuristics, not beat/downbeat, word-boundary, or section detection. Inspect the PNG, zoom and listen; never auto-delete candidates. No source or project was changed.'}
        png, js = Path(tmp)/'waveform.png', Path(tmp)/'waveform.json'
        local.run([local.external('ffmpeg'), '-v', 'error', '-nostdin', '-n', '-f', 'f32le', '-ar', str(RATE),
                   '-ac', str(channels), '-i', raw, '-filter_complex', draw_filter(report),
                   '-frames:v', '1', '-update', '1', png])
        js.write_text(json.dumps(report, indent=2)+'\n')
        # No incomplete-looking deliverable after processing failure; refuse races too.
        os.link(png, out)
        try:
            os.link(js, dest)
        except OSError:
            out.unlink()
            raise
    # Return small tool output; detailed envelope is in the JSON sidecar.
    return {k:v for k,v in report.items() if k != 'envelope'}
SKILL.md
---
name: tesseract-motion
description: Create editable motion graphics locally in Tesseract, including full-frame animated scenes, typography, diagrams, lower thirds, and overlays on supplied footage. Use for focused motion-design work; use tesseract-video for assembling or revising a complete footage edit.
---

# Design motion with Tesseract

Create a moving composition with a clear visual idea. Establish what the motion communicates, its rhythm, the destination canvas, and whether this is a full scene or an overlay that must leave footage visible.

When designing an ad opening, read [ad hooks](references/ad-hooks.md): develop a clear visual idea tied to the actual product/message, align its key arrival to confirmed opening music cues when available, and preview it together with the next scene. Sound and transitions strengthen the idea; they are not a required effects stack.

1. Resolve this skill's root from the directory containing this `SKILL.md`. Read [CLI installation](references/installation.md) and check the required Tesseract version once per session. Use the resolved executable path throughout the task. Read [local operation](references/local-operation.md) and [native authoring](references/native-authoring.md).
2. Read [motion design](references/motion-design.md). Use the supplied brand/style references. For an overlay, inspect the actual footage at the intended placement and timing; design around the subject, text, and action.
3. Choose a useful mechanism—typography, paths, masks, precomps, a diagram, or a media treatment. Use native layers and a coordinated AnimationGraph. Existing footage and graphic assets are welcome; a single prerendered image of the design defeats editable motion.
4. Retrieve the specific definitions needed from `tsrct project schema` or `tsrct project schema --document`. [FX authoring](references/fx-authoring.md) and [motion](references/motion.md) show working wire shapes and local rendering, not mandatory designs. For subject occlusion use actual person mattes when supported, not an approximate hand-drawn silhouette.
5. When supplied or previously generated music sets the rhythm, read [waveform editing](references/waveform-editing.md): open a music waveform, listen to confirm accents and phrase changes, and map animation arrivals/reveals to the selected music events before authoring timing. Preserve the chosen track; do not mistake every peak for a beat or sync every movement mechanically. Read [sound design](references/sound-design.md) when the graphic is part of an audible video. Add a local accent or transition sound when it strengthens the movement; leave it quiet when speech or the existing mix already carries the moment. Keep sound in separate Audio layers, and use [audio and timing](references/audio-and-timing.md) for importing files, placement, and envelopes. A silent-overlay request stays silent.
6. Use [filmstrip review](references/filmstrip-review.md) to inspect and correct the saved animation during construction. Finish with [review and delivery](references/review-and-delivery.md), including playback in context.

Output an editable composition in a portable `.tsrct` document and a rendered preview. Composite overlays over the supplied footage as native layers. For a transparent overlay, use `export --codec prores4444 --fx-solo compositionId:layerId --output overlay.mov` and verify the alpha channel. Solo export covers the target’s active window without audio; full-project MOV export preserves the project mix.

Prefer supported effects and native text/shape animation. For an unusual visual treatment use custom WGSL only after reading the [custom shader contract](references/custom-shader.md). Never invent API fields, keyframe formats, or a capability such as arbitrary 3D mesh import that this runtime does not provide.