SKILL DETAIL
deepline-pre-research
code.deepline.com/deepline-pre-research
The Deepline pre-research skill is used to perform a last30days-style pre-research pass before building or running a research/enrichment workflow. It helps discover critical public, private, CRM, workflow, social, and web data sources, compare provider coverage, estimate Deepline credit cost, and recommend a source plan. The skill also supports building custom language/messaging from buyer, competitor, community, and CRM evidence. The skill emphasizes live web search and exact artifact resolution, ensuring each dataset is resolved to its exact file or endpoint, with canonical URL, mirror, parser, and government statistical registry recorded. It prefers external APIs over scraping and treats private data sources as first-class. The output is a source plan, workflow design, CSV schema, play spec, or final research brief.
Installation
npx skills add https://github.com/code.deepline.com --skill deepline-pre-research
스킬 파일
SKILL.md
최근 동기화 · 2026. 8. 29.
evals/last30days-public-private-corpus.json›
{
"name": "last30days-public-private-corpus",
"version": "1.0.0-reconstructed",
"description": "20 real last30days GTM prompts used to eval the deepline-pre-research query planner. Each case asserts route classification, source-family coverage, and extraction-key emission are same-or-better than what last30days actually returned.",
"provenance": {
"reconstructedOn": "2026-08-10",
"reconstructedFrom": [
"~/.deepline-evals/eval-deepline-pre-research-corpus-fixed-1/public_private_corpus_results.json (20/20 pass)",
"~/.deepline-evals/eval-deepline-pre-research-standalone-deepline-pre-research-public-private-corpus-1/public_private_corpus_results.json (5/20 pre-fix)"
],
"note": "Original corpus JSON was never committed and is absent from disk. Case fields are recovered exactly from preserved expected_* values in the run outputs."
},
"defaults": {
"depth": "deep",
"fromDate": null,
"toDate": null,
"requiredBaseSources": [
"reddit",
"web",
"x"
],
"requiredBaseExtractionKeys": [
"company_names",
"dataset_or_api_names",
"domains",
"handles",
"linkedin_urls",
"subreddits",
"urls"
]
},
"cases": [
{
"id": "auc-propensity-lead-scoring",
"topic": "improve AUC propensity model small sample lead scoring win-loss",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"GitHub",
"Web",
"Instagram",
"Reddit",
"TikTok",
"X"
],
"mustIncludeSources": [
"github",
"instagram",
"tiktok"
],
"mustIncludeExtractionKeys": [
"pain_phrases"
],
"regression": {
"preFixRoute": "breaking_news",
"preFixMissingExtractionKeys": [
"pain_phrases"
],
"preFixSameOrBetter": false
}
},
{
"id": "paid-ads-audience-enrichment-workflow",
"topic": "B2B paid ads audience enrichment customer match workflow personal email hashes LinkedIn matched audiences Meta Google match rate",
"kind": "private_workflow",
"expectedQueryTypes": [
"private_workflow"
],
"last30daysSources": [
"Reddit",
"X",
"TikTok",
"Web"
],
"mustIncludeSources": [
"crm",
"tiktok",
"warehouse",
"workflow"
],
"mustIncludeExtractionKeys": [
"crm_object_ids",
"workflow_run_ids"
],
"regression": {
"preFixRoute": "private_workflow",
"preFixMissingExtractionKeys": [],
"preFixSameOrBetter": true
}
},
{
"id": "claude-code-gtm-workflows",
"topic": "Claude Code GTM workflows",
"kind": "private_workflow",
"expectedQueryTypes": [
"private_workflow"
],
"last30daysSources": [
"Reddit",
"X"
],
"mustIncludeSources": [
"crm",
"warehouse",
"workflow"
],
"mustIncludeExtractionKeys": [
"workflow_run_ids"
],
"regression": {
"preFixRoute": "breaking_news",
"preFixMissingExtractionKeys": [
"workflow_run_ids"
],
"preFixSameOrBetter": false
}
},
{
"id": "ai-sales-enrichment-workflows",
"topic": "AI sales enrichment workflows",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit"
],
"mustIncludeSources": [
"github"
],
"mustIncludeExtractionKeys": [
"pain_phrases"
],
"regression": {
"preFixRoute": "gtm_dataset",
"preFixMissingExtractionKeys": [],
"preFixSameOrBetter": true
}
},
{
"id": "ai-sales-enrichment-agent",
"topic": "AI sales enrichment workflows --agent",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit",
"X",
"YouTube",
"Web"
],
"mustIncludeSources": [
"youtube"
],
"mustIncludeExtractionKeys": [],
"regression": {
"preFixRoute": "gtm_dataset",
"preFixMissingExtractionKeys": [],
"preFixSameOrBetter": true
}
},
{
"id": "claude-code-gtm-agent",
"topic": "Claude Code GTM workflows --agent",
"kind": "private_workflow",
"expectedQueryTypes": [
"private_workflow"
],
"last30daysSources": [
"Reddit",
"X"
],
"mustIncludeSources": [
"crm",
"warehouse",
"workflow"
],
"mustIncludeExtractionKeys": [
"workflow_run_ids"
],
"regression": {
"preFixRoute": "breaking_news",
"preFixMissingExtractionKeys": [
"workflow_run_ids"
],
"preFixSameOrBetter": false
}
},
{
"id": "plg-signup-onboarding-retention",
"topic": "PLG signup onboarding retention developer tools CLI-first signup flow activation",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit",
"X",
"TikTok",
"Instagram",
"Hacker News",
"Web"
],
"mustIncludeSources": [
"hn",
"instagram",
"tiktok",
"youtube"
],
"mustIncludeExtractionKeys": [
"category_terms",
"pain_phrases"
],
"regression": {
"preFixRoute": "breaking_news",
"preFixMissingExtractionKeys": [
"category_terms",
"pain_phrases"
],
"preFixSameOrBetter": false
}
},
{
"id": "salesforce-opportunity-win-loss",
"topic": "Salesforce opportunity fields and signals for predictive win-loss lead scoring B2B",
"kind": "private_workflow",
"expectedQueryTypes": [
"private_workflow"
],
"last30daysSources": [
"Reddit",
"X",
"TikTok",
"Instagram",
"Web"
],
"mustIncludeSources": [
"crm",
"instagram",
"tiktok",
"warehouse",
"workflow"
],
"mustIncludeExtractionKeys": [
"crm_object_ids",
"deal_or_opportunity_ids"
],
"regression": {
"preFixRoute": "private_workflow",
"preFixMissingExtractionKeys": [],
"preFixSameOrBetter": true
}
},
{
"id": "seo-page-level-fixes",
"topic": "best page-level SEO fixes for internal linking canonical tags freshness signals noindex pruning",
"kind": "public_research",
"expectedQueryTypes": [
"concept",
"how_to"
],
"last30daysSources": [
"Reddit",
"X",
"TikTok",
"Web"
],
"mustIncludeSources": [
"hn"
],
"mustIncludeExtractionKeys": [],
"regression": {
"preFixRoute": "breaking_news",
"preFixMissingExtractionKeys": [],
"preFixSameOrBetter": false
}
},
{
"id": "azure-warehouse-salesforce-reverse-etl",
"topic": "moving Azure data warehouse data into Salesforce reverse ETL and zero copy options",
"kind": "private_workflow",
"expectedQueryTypes": [
"private_workflow"
],
"last30daysSources": [
"Reddit",
"X",
"YouTube",
"TikTok",
"Instagram",
"Web"
],
"mustIncludeSources": [
"crm",
"instagram",
"tiktok",
"warehouse",
"workflow",
"youtube"
],
"mustIncludeExtractionKeys": [
"crm_object_ids",
"workflow_run_ids"
],
"regression": {
"preFixRoute": "custom_language",
"preFixMissingExtractionKeys": [],
"preFixSameOrBetter": false
}
},
{
"id": "blitzapi-signaliz-providers",
"topic": "BlitzAPI and Signaliz B2B data sourcing providers",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit",
"X",
"TikTok",
"Web"
],
"mustIncludeSources": [
"github",
"tiktok"
],
"mustIncludeExtractionKeys": [
"competitor_names"
],
"regression": {
"preFixRoute": "breaking_news",
"preFixMissingExtractionKeys": [
"competitor_names"
],
"preFixSameOrBetter": false
}
},
{
"id": "data-broker-reseller-compliance",
"topic": "data broker reseller compliance best practices misuse prevention FCRA GLBA DPPA acceptable use",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit",
"X",
"TikTok",
"Web"
],
"mustIncludeSources": [
"github",
"tiktok"
],
"mustIncludeExtractionKeys": [
"category_terms"
],
"regression": {
"preFixRoute": "how_to",
"preFixMissingExtractionKeys": [
"category_terms"
],
"preFixSameOrBetter": false
}
},
{
"id": "partner-enablement-kit",
"topic": "partner enablement kit best practices for B2B SaaS partner programs",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit",
"X",
"YouTube",
"TikTok",
"Instagram",
"Web"
],
"mustIncludeSources": [
"instagram",
"tiktok",
"youtube"
],
"mustIncludeExtractionKeys": [
"category_terms",
"pain_phrases"
],
"regression": {
"preFixRoute": "how_to",
"preFixMissingExtractionKeys": [
"category_terms",
"pain_phrases"
],
"preFixSameOrBetter": false
}
},
{
"id": "plg-retention-levers",
"topic": "PLG retention levers sales-led usage retention signup funnel conversion",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit",
"X",
"TikTok",
"Instagram"
],
"mustIncludeSources": [
"instagram",
"tiktok"
],
"mustIncludeExtractionKeys": [
"pain_phrases"
],
"regression": {
"preFixRoute": "breaking_news",
"preFixMissingExtractionKeys": [
"pain_phrases"
],
"preFixSameOrBetter": false
}
},
{
"id": "identity-fraud-kyc-b2b",
"topic": "identity verification fraud prevention account takeover data breach KYC regulation B2B fintech banking",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit",
"X",
"TikTok",
"Instagram",
"Polymarket",
"Web"
],
"mustIncludeSources": [
"instagram",
"polymarket",
"tiktok"
],
"mustIncludeExtractionKeys": [
"category_terms"
],
"regression": {
"preFixRoute": "breaking_news",
"preFixMissingExtractionKeys": [
"category_terms"
],
"preFixSameOrBetter": false
}
},
{
"id": "production-account-scoring",
"topic": "production B2B propensity and rules-based account scoring benchmarks and validation",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit",
"X",
"YouTube",
"TikTok",
"Instagram",
"Web"
],
"mustIncludeSources": [
"instagram",
"tiktok",
"youtube"
],
"mustIncludeExtractionKeys": [
"pain_phrases"
],
"regression": {
"preFixRoute": "breaking_news",
"preFixMissingExtractionKeys": [
"pain_phrases"
],
"preFixSameOrBetter": false
}
},
{
"id": "event-follow-up-sequence",
"topic": "in-person B2B event follow-up best practices email LinkedIn sequence post event",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit",
"X",
"YouTube",
"TikTok",
"Instagram"
],
"mustIncludeSources": [
"instagram",
"tiktok",
"youtube"
],
"mustIncludeExtractionKeys": [
"pain_phrases",
"personas"
],
"regression": {
"preFixRoute": "how_to",
"preFixMissingExtractionKeys": [
"pain_phrases",
"personas"
],
"preFixSameOrBetter": false
}
},
{
"id": "paid-ads-match-rates",
"topic": "B2B paid ads audience enrichment match rates Google Customer Match Meta Custom Audiences LinkedIn Matched Audiences personal email hashes LinkedIn URLs",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit",
"X",
"TikTok",
"Instagram",
"Web"
],
"mustIncludeSources": [
"instagram",
"tiktok"
],
"mustIncludeExtractionKeys": [],
"regression": {
"preFixRoute": "gtm_dataset",
"preFixMissingExtractionKeys": [],
"preFixSameOrBetter": true
}
},
{
"id": "plg-activation-churn-signals",
"topic": "PLG activation retention churn signals B2B SaaS developer tools onboarding support bugs billing overcharges",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit",
"X",
"YouTube",
"TikTok",
"Hacker News",
"Web"
],
"mustIncludeSources": [
"hn",
"tiktok",
"youtube"
],
"mustIncludeExtractionKeys": [
"category_terms",
"pain_phrases"
],
"regression": {
"preFixRoute": "breaking_news",
"preFixMissingExtractionKeys": [
"category_terms",
"pain_phrases"
],
"preFixSameOrBetter": false
}
},
{
"id": "new-product-launch-account-targeting",
"topic": "how companies target accounts for new product launches with AI tools or historical scoring models",
"kind": "public_gtm_dataset",
"expectedQueryTypes": [
"gtm_dataset"
],
"last30daysSources": [
"Reddit",
"X",
"YouTube",
"TikTok",
"Instagram",
"Web"
],
"mustIncludeSources": [
"instagram",
"tiktok",
"youtube"
],
"mustIncludeExtractionKeys": [
"pain_phrases"
],
"regression": {
"preFixRoute": "breaking_news",
"preFixMissingExtractionKeys": [
"pain_phrases"
],
"preFixSameOrBetter": false
}
}
]
}
references/fanout-consolidation.md›
# Fanout And Consolidation
Use this reference when comparing `/deepline-pre-research` to `last30days` or when designing the actual Deepline workflow.
## Baseline To Preserve
`last30days` gets its usefulness from a two-phase retrieval loop:
1. Broad fanout across public sources in parallel.
2. Supplemental fanout from discovered handles, subreddits, entities, and source-specific leads.
3. Normalization into comparable result fields.
4. Relevance, recency, engagement, and source-quality scoring.
5. Same-source dedupe plus cross-source linking.
6. Source stats, errors, missing-source nudges, and synthesis.
Deepline should preserve that shape. Pre-research itself should prioritize public-source discovery and a `last30days`-style "What I learned" synthesis for the GTM problem. Private and workflow data are identified as join targets and handed to `deepline-gtm`, `deepline-analytics`, or a play after the public map is clear.
## Deepline Fanout Contract
### Stage 0: Parse
Extract objective, entity scope, time window, public sources, private sources, dataset leads, custom-language outputs, and final artifact. Turn unclear scope into explicit assumptions before any provider calls.
### Stage 1: Broad Parallel Fanout
Search in parallel across public sources first:
- social/community: Reddit threads, Reddit comments, X/Twitter, YouTube, TikTok, Instagram, HN, Polymarket, Bluesky, Truth Social
- web/source discovery: web, news, docs, blogs, GitHub, directories, app stores, reviews
- public and niche datasets: public records, registries, licenses, inspections, permits, government datasets, open CSVs/APIs, professional directories, association/member lists, accreditation databases
- GTM public signals: company/account, people/contact, jobs, hiring, technographics, funding, review volume, ads and public activity
- Deepline provider catalog is not part of Stage 1. Search and describe Deepline tools only after the public-source synthesis identifies which source families, private joins, probes, costs, and gaps matter.
Do not treat private datasets as the broad fanout unless the user specifically asks for proprietary-data analysis. For GTM work, identify private/proprietary datasets as Stage 7 join targets after public-source discovery.
### Stage 2: Supplemental Fanout
For every useful result, extract follow-up keys and run targeted searches:
- subreddits, handles, channels, creators, domains, URLs, GitHub repos
- company domains, LinkedIn URLs, account ids, CRM object ids, deal/opportunity ids
- workflow/play/run ids, dataset ids, sheet ids, warehouse dimensions
- public-record registry names, license ids, permit ids, agency names
- healthcare registry examples such as NPI registry/provider taxonomy; these are generic-route public sources when no native catalog tool exists
- persona terms, pain phrases, objections, competitor names, category language
Supplemental searches should be attributable to the lead that produced them so evidence clusters remain explainable.
### Stage 3: Normalize
Every result should fit a common evidence shape:
| Field | Meaning |
| --- | --- |
| `source_family` | social, web, company, person, jobs, technographic, funding, CRM, warehouse, workflow, support, sheet, custom-language |
| `route_status` | native, generic route, private connector, gap |
| `tool_or_provider` | Deepline tool id, provider, connector, or proposed gap |
| `source_id_or_url` | stable id, URL, CRM id, workflow id, or dataset id |
| `title_or_label` | human-readable result label |
| `text_or_excerpt` | quoted evidence or normalized text |
| `author_or_owner` | author, channel, CRM owner, source system, or blank |
| `timestamp` | created/published/observed date when available |
| `engagement_or_outcome` | public engagement or private outcome metric |
| `join_key` | domain, email, LinkedIn URL, account id, contact id, dataset key |
| `cost_basis` | Deepline credit basis or unknown until describe/probe |
| `confidence` | extraction/match confidence |
| `provenance` | query, supplemental lead, run id, file id, or connector object path |
### Stage 4: Score
Preserve `last30days`-style relevance, recency, engagement, and source-quality scoring. Add Deepline-specific score inputs:
- materializability: can this become a repeatable dataset/API/provider call?
- private-outcome value: does it connect to conversion, retention, revenue, support load, or usage?
- activation value: can the result drive a CRM field, enrichment column, campaign step, or play branch?
- join confidence: are stable join keys present?
- cost/coverage value: does the source justify its Deepline credit cost?
- custom-language fit: is the phrase tied to the right persona, segment, stage, and context?
### Stage 5: Dedupe And Cluster
Use at least the `last30days` dedupe discipline:
- canonical URL/source id match
- normalized text similarity
- token or n-gram similarity
- per-source dedupe before cross-source linking
Then add Deepline-specific consolidation:
- company identity joins by domain, CRM account id, LinkedIn company URL, or verified provider id
- person identity joins by email, LinkedIn profile URL, CRM contact id, and company/title
- dataset canonicalization by agency/API/repo URL and stable dataset id
- private/public evidence clusters that keep CRM/workflow ids separate from public citations
- custom-language clusters grouped by phrase, persona, pain, objection, and reuse field
### Stage 6: Coverage Nudge
Before synthesis, emit a coverage table for each required source family:
- `searched and useful`
- `searched and weak`
- `searched and errored`
- `not searched because missing credentials`
- `not available in catalog`
- `not relevant to this job`
Only synthesize after the user can see what is strong, weak, missing, and cost-unknown.
### Stage 7: Segment Synthesis And Handoff
Return a concise GTM research report before the implementation plan:
1. "Research Report" with key findings and a "What I learned" section, following `last30days --agent` style.
2. "Best public sources" - public registries, communities, reviews, directories, datasets, and source leads with route status and join keys.
3. "What to join later" - proprietary CRM, warehouse, workflow, support, customer CSV, product, or outcome datasets that would prove the public signals matter.
4. "Deepline route" - only now map sources to Deepline tools, costs, probes, and gaps.
For example, in healthcare provider analysis, "NPI registry has no native Deepline tool" is not the conclusion. The conclusion is: "NPI taxonomy is a high-recall seed but over-includes non-ICP storefronts; pull it via generic web/API route, then validate against Maps/website identity and join proprietary win/loss or CRM outcomes."
## Better-Than-last30days Criteria
Claim Deepline is better only when the workflow:
- covers the same public/community sources or explicitly marks gaps
- identifies private datasets and customer-owned evidence to join after public discovery
- quotes Deepline-facing cost before scaled execution
- produces joinable rows, not just prose
- preserves citations, raw paths, run ids, and connector object ids
- outputs a workflow/play shape that can be rerun
- separates exact market language from rewritten copy
references/last30days-gtm-corpus.md›
# Last30Days GTM Corpus
This summarizes the relevant saved `last30days` runs under `~/Documents/Last30Days` as of this review. Use it as source-selection memory for `/deepline-pre-research`; do not depend on `last30days` at runtime.
## Corpus Shape
Filename classification found:
| Family | Approx. relevant runs | What it means for Deepline |
| --- | ---: | --- |
| Provider strategy / enrichment / waterfall | 144 | Provider comparison, cost, coverage, waterfall design, and data-quality runs are the most common GTM use case. |
| Contact enrichment / identity / LinkedIn / email / phone | 137 | Name/company to LinkedIn/email/phone, same-name disambiguation, small-company misses, and verification are core. |
| GTM content/community/social/custom language | 119 | LinkedIn/X/YouTube/community sources are useful for messaging, hooks, voice, event/community discovery, custom language, objection phrasing, and market language. |
| CRM/workflow/private data | 83 | Salesforce, HubSpot, Snowflake, Marketo, workflow runs, PLG/product usage, and call/email/Slack transcripts must be first-class. |
| Signals/scoring/intent | 72 | Buying triggers, propensity, launch/cross-sell, audience matching, and account prioritization need mixed private + public evidence. |
| Public records / niche datasets | 51 | SMB, restaurant, grocery, nursing-home, construction, legal-entity, permit/license, event-attendee, and owner lookup runs are high-signal. |
The broad lesson: GTM value came less from "what are people saying?" and more from "what public or private dataset did the discussion point us toward?"
## Relevant Run Families
### 1. Provider Strategy And Waterfalls
Representative runs:
- `best-gtm-data-sources-smb-consumer-services-companies-raw-v3.md`
- `b2b-data-enrichment-pricing-quality-benchmark-provider-comparison-apollo-zoominfo-clay-cognism-lusha-peopledatalabs-international-coverage-accuracy-raw.md`
- `clay-ads-primer-metadata-io-best-providers-best-practices-b2b-audience-personal-email-phone-enrichment-meta-google-linkedin-raw.md`
- `clay-waterfall-enrichment-credits-raw-v3.md`
- `gtm-tool-pricing-stack-migration-claude-code-cli-pain-language-raw-v3.md`
Reusable source pattern:
- Start with community evidence from Reddit/X for pain, coverage complaints, and current vendor sentiment.
- Use web/docs only to verify vendor claims and pricing.
- Translate into Deepline provider routing: direct provider, waterfall, or gap.
- Always distinguish "lead source" from "contact enrichment" from "verification"; many bad workflows blur these.
Implication for `/deepline-pre-research`: provider strategy output must include coverage basis, cost basis, expected miss patterns, and fallback order.
### 2. Contact Enrichment And Identity Resolution
Representative runs:
- `linkedin-profile-url-lookup-from-name-api-scraping-resolution-raw-v3.md`
- `linkedin-url-lookup-pipelines-from-name-and-company-providers-validation-false-positives-gtm-engineering-raw-v3.md`
- `finding-people-not-on-linkedin-collections-specialists-admin-roles-lower-level-prospecting-raw-v3.md`
- `finding-contacts-at-small-companies-under-50-employees-when-data-providers-return-zero-results-gtm-engineer-sales-ops-raw-v3.md`
- `email-pattern-guessing-false-positive-wrong-person-verification-identity-confirmation-b2b-cold-outreach-raw-v3.md`
- `best-b2b-mobile-phone-number-data-providers-accuracy-compari-raw.md`
Reusable source pattern:
- LinkedIn search/scrape is often the identity anchor, but not sufficient.
- Small companies and non-US targets frequently miss in large B2B databases; search/web/registry fallback matters.
- Same-name and stale-role failures are common enough to be a first-class gate.
- Email guessing without identity validation creates wrong-person risk.
Implication: pre-research must require identity evidence fields, not just returned contact fields: current company, current title, work-history corroboration, source URL, and validation status.
### 3. Public Records And Niche Datasets
Representative runs:
- `restaurant-owner-phone-number-data-sources-public-records-sos-liquor-license-health-permit-business-license-api-2025-2026-raw-v3.md`
- `public-data-sources-independent-grocery-store-owners-nyc-raw-v3.md`
- `cms-nursing-home-public-data-deficiencies-staffing-star-rati-raw.md`
- `building-permit-and-construction-project-data-providers-raw-v3.md`
- `cannabis-public-data-sources-license-registry-regulatory-filings-raw-v3.md`
- `legal-entity-data-gtm-corporate-hierarchy-parent-subsidiary-api-raw-v3.md`
- `event-attendee-list-conference-prospecting-raw-v3.md`
Reusable source pattern:
- The strongest GTM datasets are often vertical public records: permits, licenses, inspections, deficiencies, business registries, county records, public ownership, and event attendance.
- Social/community results are useful because they reveal that a dataset exists, not because the posts are final evidence.
- Government/open-data datasets need fetch/download/normalize/join steps before enrichment.
Implication: `/deepline-pre-research` should promote "dataset lead discovery" as a source family and ask whether a public registry can be materialized before defaulting to generic company databases.
### 4. Signals, Scoring, And Propensity
Representative runs:
- `buying-signals-sales-prospecting-hard-to-find-rare-unique-intent-data-raw-hard-signals.md`
- `ai-propensity-to-buy-and-external-signals-for-b2b-saas-expansion-raw-pcc.md`
- `default-account-scoring-gtm-workflows-lead-routing-ai-raw.md`
- `how-companies-target-accounts-for-new-product-launches-with-ai-tools-or-historical-scoring-models-raw-launch.md`
- `backtesting-prompts-for-gtm-lead-scoring-and-qualification-apply-scoring-framework-to-last-100-leads-validate-model-against-historical-data-raw-v3.md`
Reusable source pattern:
- Public buying signals are useful when they can be tied to a dated event: hiring, funding, launches, job posts, regulatory changes, intent-like behavior, ads/audience membership, or product usage.
- Private data decides whether the signal is predictive: won/lost accounts, conversion, pipeline, usage, and expansion history.
- Backtesting needs historical CRM/warehouse data, not just a prompt.
Implication: source plans for scoring must include both public signal retrieval and private outcome labels, with join keys and a validation split/backtest plan.
### 5. CRM, Workflow, Warehouse, And Product Usage
Representative runs:
- `ai-agent-workflow-salesforce-separate-database-scoring-inbou-raw.md`
- `data-warehouse-gtm-engineering-plg-product-usage-workflows-raw.md`
- `b2b-sales-signal-workflows-snowflake-listagg-limits-cost-attribution-per-step-llm-cost-caps-crustdata-enrichment-middleware-traps-brief-format-raw-session-retro.md`
- `marketo-cross-sell-campaign-orchestration-using-propensity-scores-and-account-expansion-signals-raw.md`
- `hubspot-button-webhook-enrich-record-external-api-raw-v3.md`
- `ai-automatically-update-todo-prioritization-list-email-slack-call-transcripts-task-ownership-raw-v3.md`
Reusable source pattern:
- CRM should not be only an output destination. It is a private evidence source: lifecycle, owner, stage, activity, field history, source attribution.
- Warehouse/product usage/workflow runs are often the real truth layer for PLG and expansion.
- Call transcripts, Slack/email, and workflow outputs are useful when converted into dated, attributable signals.
Implication: `/deepline-pre-research` should ask for CRM/warehouse/workflow access before finalizing any GTM source plan.
### 6. Custom Language, Messaging, And Community Language
Representative runs:
- `cold-email-outreach-hooks-personalization-what-s-working-2026-raw-v3.md`
- `linkedin-post-hooks-viral-2026-raw-v3.md`
- `founder-led-gtm-linkedin-content-claude-code-ai-gtm-engineering-raw-v3.md`
- `b2b-gtm-lead-magnets-that-convert-2026-lead-magnet-ideas-for-revops-and-growth-engineers-ai-gtm-data-tooling-lead-magnets-coreyhaines-lead-magnet-raw-v3.md`
- `youtube-seo-event-video-clips-thumbnails-b2b-saas-ranking-2026-raw.md`
- `gtm-tool-pricing-stack-migration-claude-code-cli-pain-language-raw-v3.md`
- `warm-outbound-visitor-message-template-vs-ai-personalization-raw-v3.md`
- `post-call-follow-up-email-b2b-saas-what-to-include-raw-v3.md`
Reusable source pattern:
- X/LinkedIn/Reddit/YouTube are strongest for language, objections, proof points, hooks, and examples.
- Sales calls, reply emails, CRM notes, support threads, and win/loss notes are the strongest private-language sources because they preserve the buyer's exact words and outcome context.
- These sources are weaker for durable facts unless corroborated by docs, product pages, official sources, or datasets.
- Engagement counts help rank messaging patterns, but should not become factual proof.
Implication: split "market language" from "verified dataset/source" in the output, and produce structured language assets when requested: pain phrases, objection language, category terms, hooks, subject lines, opener snippets, CTA language, and account/persona-specific personalization fields.
Custom language workflow:
1. Gather exact phrases from public community and private customer sources.
2. Cluster them by persona, pain, objection, desired outcome, and maturity stage.
3. Preserve representative verbatim snippets with source context.
4. Rewrite into target channel formats only after the evidence table exists.
5. Return both evidence columns and generated copy columns so the workflow can be rerun.
## Mining-Safety Example
A Codex session captured the strongest GTM lesson:
- `/last30days` found an open government mining-safety dataset via a Reddit post.
- It found a predictive mine-safety research paper via X.
- It found a government-agency post pointing to roughly 3M violation records.
- The agent built a risk model, scored 6,537 mines, found safety decision makers, validated emails/LinkedIn URLs, and wrote Lemlist campaigns.
This should become the canonical `/deepline-pre-research` mental model:
1. Use social/community/web to discover the dataset and modeling literature.
2. Materialize the dataset.
3. Build the scoring or filtering layer.
4. Enrich account/contact records.
5. Activate into outbound/workflow tools.
The social post is not the deliverable. The operational dataset and workflow are the deliverable.
## Source Families To Prioritize In Deepline
High-priority because repeated GTM logs used them:
- Reddit comments/threads: pain language, niche datasets, practitioner failures.
- X/Twitter: current practitioner tactics, dataset/paper pointers, named operators.
- YouTube/TikTok/Instagram transcripts and captions: phrasing, examples, demos, creator framing, objections, and hooks.
- Web/news/docs/GitHub: official docs, repos, public data inventories, pricing pages, API docs.
- ScrapeCreators-style social data: Reddit comments plus TikTok/Instagram captions are useful enough to justify native support or a strong generic provider route.
- Apify actors: pragmatic route for LinkedIn, social scraping, and odd public web datasets when native providers are missing.
- Public records/government/open-data: vertical-specific GTM alpha.
- CRM/warehouse/product/workflow data: private truth layer for scoring and activation.
- Contact/enrichment/verification stack: Apollo, Dropleads, Hunter, LeadMagic, BetterContact, Icypeas, RocketReach, ContactOut, Wiza, PDL, ZeroBounce-style validation, and LinkedIn scraping.
- Campaign activation: Lemlist, Smartlead, Instantly, HeyReach, Marketo, HubSpot/Salesforce actions where relevant.
## Add To Every GTM Pre-Research Plan
Before choosing providers, answer:
1. What public/community source can reveal hidden datasets or unusual signals?
2. What private dataset proves the signal matters?
3. What registry/API/source can be materialized into rows?
4. What contact/enrichment path reaches the right buyer after scoring?
5. What activation surface receives the result?
6. What cost/coverage probe validates the plan before scale?
7. What exact buyer/community/customer language should be preserved for messaging or personalization?
references/query-design.md›
# Query Design
`/deepline-pre-research` should not send the user's raw prompt to every source. The query-design layer is the core behavior to preserve from `last30days`.
Attribution: the query-cleaning, query-type, source-tiering, and supplemental extraction patterns are adapted from `mvanhorn/last30days-skill` under the MIT License. See `../THIRD_PARTY_NOTICES.md`.
Use `scripts/query_design.py` to generate the query plan:
```bash
python3 .skills/deepline-pre-research/scripts/query_design.py "best GTM data sources for SMB consumer services companies" --depth deep
```
In Deepline runtime, use the native API route only after the public-source fanout has produced a first synthesis and source map. The API translates findings into Deepline routes and costs; it must not replace the research pass:
```bash
curl -s "$DEEPLINE_API_BASE_URL/api/v2/pre-research/plan" \
-H "Authorization: Bearer $DEEPLINE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"objective":"best GTM data sources for SMB consumer services companies","depth":"deep"}'
```
The script emits JSON with:
- `query_type`: product, concept, opinion, how_to, comparison, breaking_news, prediction, gtm_dataset, private_workflow, or custom_language
- `core_subject`: a cleaned version of the prompt with question/meta/noise words removed
- `enabled_sources`: tiered source defaults for the detected query type
- `variants`: source-specific broad and fallback queries
- `supplemental_templates`: phase-two templates for discovered handles, subreddits, domains, accounts, CRM ids, datasets, and language phrases
- `extraction_keys`: what to pull from phase-one results before supplemental fanout
- `scoring_notes`: how to rank results before synthesis
## What Was Ported
Preserve these behaviors:
- Strip common prompt wrappers and research/noise words before platform searches.
- Keep different core subjects by platform. X should be aggressive because keyword search is literal; YouTube/TikTok/Instagram should keep content-type terms like tutorial, review, and tips.
- Detect query type before choosing source defaults.
- Expand Reddit into core, original, review/opinion, and problem/issue variants depending on depth and query type.
- Expand X into the core query with a recency filter, quoted compound-term OR queries, shorter keyword fallback, and strongest-token fallback.
- Run video/social caption searches separately from web searches.
- Treat private/workflow/custom-language queries as first-class query types, not post-processing.
- Extract phase-one handles, hashtags, subreddits, domains, datasets, CRM ids, workflow ids, personas, pain phrases, objections, competitors, and category terms for supplemental fanout.
## Deepline Additions
The Deepline query planner extends `last30days` in four ways:
- `gtm_dataset` query type for provider strategy, public records, account data, contact data, intent data, and enrichment waterfalls.
- `private_workflow` query type for CRM, warehouse, product usage, workflow runs, and RevOps evidence.
- `custom_language` query type for buyer language, objections, competitor/category phrasing, and campaign hooks.
- Provider-catalog and cost-awareness hooks so synthesized research findings can become Deepline source plans with approval gates after the public pass.
## Implementation Contract
Any Deepline runtime implementation should follow this order:
1. Build query plan with `query_design.py` logic or equivalent TypeScript port.
2. Execute public broad variants across the selected source families.
3. Extract supplemental keys from phase-one evidence.
4. Execute targeted supplemental variants.
5. Normalize every result to the evidence schema in `fanout-consolidation.md`.
6. Score, dedupe, cluster, and emit coverage nudges.
7. Synthesize the `last30days`-style GTM research report.
8. Search/describe Deepline tools for the useful source families, private joins, costs, probes, and provider gaps.
Do not collapse this into one generic web search. The source-specific query variants are the quality lever.
references/source-map.md›
# Source Map
Use this reference to choose source families for `/deepline-pre-research`. Tool ids drift; always confirm with `deepline tools search` and `deepline tools describe`.
## Required Dataset Inventory
To replicate `last30days`-style functionality inside Deepline, every pre-research plan must account for these dataset families. Some are already first-class Deepline providers, some can be reached through generic routes, and some are explicit integration gaps to add.
| Dataset family | Required evidence fields | Deepline route today | Parity status |
| --- | --- | --- | --- |
| Reddit threads | post URL, subreddit, title/body, author when available, score/upvotes, comments count, created date | public-source discovery first; translate to catalog route after synthesis | generic route or gap |
| Reddit comments | comment text, comment URL/permalink, author when available, upvotes, parent post, created date | ScrapeCreators-style provider or vetted `apify` actor | gap unless catalog finds native support |
| X/Twitter posts | post URL/id, handle, text, timestamp, likes, reposts, replies, quote count, optional entity handle search | public-source discovery first; translate to catalog route after synthesis | generic route or gap |
| YouTube search/transcripts | video URL/id, channel, title, publish date, views/likes, transcript excerpts | public-source discovery first; translate to catalog route after synthesis | generic route or gap |
| TikTok | video URL/id, creator, caption/transcript, hashtags, timestamp, views/likes/comments | ScrapeCreators-style provider or `apify` actor | gap or generic route |
| Instagram Reels/posts | URL/id, creator, caption/transcript, timestamp, views/likes/comments | ScrapeCreators-style provider or `apify` actor | gap or generic route |
| Hacker News | story/comment URL, title/text, author, points, comments, created date | Algolia HN native provider recommended; web/search fallback | gap or generic route |
| Polymarket | market URL/id, question, outcomes, odds, movement, volume/liquidity, close date | Gamma API native provider recommended; web/search fallback | gap or generic route |
| Bluesky | post URL/id, handle, text, timestamp, likes/reposts/replies | native BYOK/app-password route recommended | gap unless catalog finds support |
| Truth Social | post URL/id, handle, text, timestamp, likes/reposts/replies | native/BYOK or vetted extraction route recommended | gap unless catalog finds support |
| Web/news/blogs/docs | URL, title, snippet/full text, publish date, source domain, citations | `serper`, `exa`, `parallel`, `firecrawl`, `deeplineagent` | native/generic route |
| Public registries and niche public datasets | registry/API URL, dataset title, row identifiers, entity names, addresses, status/taxonomy/license fields, update date, source agency | generic web/API discovery through `serper`, `exa`, `parallel`, `firecrawl`, `generic_http`, `deeplineagent`; native tool only when catalog finds one | generic route unless catalog finds native support |
| Company/account datasets | domain, name, LinkedIn URL, firmographics, funding, industry, geo, employee count, source URL | `crustdata`, `apollo`, `dropleads`, `openmart`, `aviato`, `peopledatalabs`, `forager` | native |
| People/contact datasets | full name, title, company, LinkedIn URL, email/phone when approved, confidence, source | `apollo`, `dropleads`, `hunter`, `leadmagic`, `bettercontact`, `icypeas`, `rocketreach`, `contactout`, `wiza`, `peopledatalabs` | native |
| Jobs/hiring signals | job title, company, location, posted date, description, URL | `crustdata`, `predictleads`, `bloomberry`, `serper`, `firecrawl` | native |
| Technographics/install base | domain/company, vendor/tool, category, detected date, confidence/source | `builtwith`, `theirstack`, `bloomberry` | native |
| Funding/company events/news | company, event type, date, amount/round when available, source URL | `crustdata`, `aviato`, `predictleads`, `serper`, `exa` | native/generic route |
| CRM data | account/contact/deal ids, owner, stage, lifecycle timestamps, activity fields, custom fields | `salesforce`, `hubspot`, `attio` | private connector |
| Warehouse/semantic metrics | metric name, dimensions, filters, result rows, rendered SQL/audit trail | `snowflake_get_semantic_layer`, `snowflake_run_semantic_query` | private connector |
| Product/workflow usage | org/user/account ids, event/run ids, timestamps, status, outputs, credits/run metadata | `plays`, workflow/session/usage tools, warehouse/customer DB | private connector |
| Support/call/docs data | transcript/note URL/id, speaker/contact, timestamp, summary/snippet, related account | CRM, Attio, Slack/docs connectors, call transcript tools where configured | private connector |
| Sheets/CSVs/customer-owned lists | row ids, domains/emails, source file/sheet id, provenance columns | user CSV, Google Sheets, Deepline Playground | private/customer source |
| Custom language corpus | exact quote/snippet, speaker/source, audience/persona, context, timestamp, engagement or CRM outcome when available | Reddit/X/LinkedIn/YouTube/TikTok/Instagram, sales calls, support tickets, CRM notes, win/loss notes, Gong/Fireflies-like transcripts, docs/sheets | native/generic/private mix |
Minimum useful parity for a last-30-days community research run is: web/news, Reddit threads, Reddit comments, X posts, YouTube transcripts, TikTok or Instagram when relevant, HN, Polymarket when relevant, and at least one private/customer context source when the job is GTM/customer-facing.
For custom language runs, minimum useful parity is: at least one community/social source, at least one long-form context source (YouTube transcript, podcast, blog, call transcript, support thread, or Reddit comments), and one business-context source such as CRM stage/outcome, persona, segment, deal notes, or target account list.
## What To Preserve From last30days
The useful pattern is not the local script itself. Preserve:
- broad source taxonomy: Reddit, X/Twitter, YouTube, TikTok, Instagram, HN, Polymarket, Bluesky, Truth Social, and web
- recency-first retrieval for trends and news
- community evidence with engagement stats
- source coverage stats and explicit missing-source nudges
- comparison mode that researches each side separately plus direct comparison
- synthesis weighted by source quality, engagement, and cross-platform agreement
- two-phase fanout: broad parallel source search, then supplemental searches from discovered handles, subreddits, domains, datasets, and entities
- common evidence rows before scoring, dedupe, clustering, and synthesis
Deepline adaptation:
- route through Deepline tools, plays, workflows, CRM connectors, and warehouse connectors
- quote Deepline credits only
- save outputs in CSVs/runs/playgrounds where agents and users can inspect them
- convert the source plan into a repeatable workflow when the user wants automation
- add private/source-of-truth joins, provider-cost checks, approval gates, and workflow activation paths that `last30days` does not own
## Public Source Families
| Need | Start With | Notes |
| --- | --- | --- |
| General web/news/source discovery | `serper`, `exa`, `parallel`, `deeplineagent`, `firecrawl` | Use search to discover URLs; use extraction for known pages or JS-rendered sources. |
| Public registries / niche datasets | `serper`, `exa`, `parallel`, `firecrawl`, `generic_http`, `deeplineagent` | Treat public registries as materializable sources even without native Deepline tools. Example: NPI registry/provider taxonomy data can be found and pulled through generic web/API search/extraction, then joined by NPI, organization name, address, phone, and taxonomy. |
| Reddit threads/comments | Discover public/community evidence first; search Deepline catalog after synthesis to pick the execution route | If full Reddit comments are core, recommend a native ScrapeCreators-style provider or a vetted Apify actor. |
| X/Twitter posts | Discover public/social evidence first; search catalog after synthesis for X/Twitter/social execution | Browser cookies are not an acceptable backend Deepline integration pattern. Prefer official/BYOK or managed provider integration. |
| YouTube search/transcripts | Discover video/transcript evidence first; search catalog after synthesis for transcript execution | Preserve channel, URL, views, publish date, and transcript excerpts. |
| TikTok/Instagram | Discover short-form/social evidence first; search catalog after synthesis for execution route | Use captions/transcripts plus engagement. For local-business contact workflows, also consider Instagram profile bio links/contact fields as candidate contact-data signals. Search ScrapeCreators unfiltered because profile tools may be categorized as `admin`, not `research`. Avoid unauthenticated brittle scraping at scale. |
| Facebook pages/profiles | Discover public page/profile contact details when local businesses, restaurants, storefronts, or social-first companies may publish email/phone/website there | Candidate route through ScrapeCreators profile tools or generic web extraction. Search ScrapeCreators unfiltered because Facebook profile tools may be categorized as `admin`, not `research`. Preserve profile URL and identity evidence. |
| Hacker News | Search/web fallback may be enough; native Algolia HN would be cheap to add | Capture points, comments, author, URL, and story age. |
| Polymarket | Search/web fallback may be enough; native Gamma API would be cheap to add | Capture odds, movement, market close date, and volume/liquidity when available. |
| Bluesky/Truth Social | Treat as explicit gap unless catalog search finds a current route | Prefer API-backed access over brittle page scraping. |
| Jobs/hiring signals | `crustdata`, `predictleads`, `bloomberry`, `theirstack`, `serper`, `firecrawl` | Jobs are high-intent and often better than marketing pages. |
| Tech stack/install base | `builtwith`, `theirstack`, `bloomberry`, `wappalyzer-like catalog hits` | Confirm whether the tool supports domain lookup, vendor lookup, or customer list search. |
| Funding/company events | `crustdata`, `aviato`, `apollo`, `predictleads`, `serper`, `exa` | Verify event date, source URL, and company domain. |
| Ads/audience signals | `dataforseo`, `google-ads-audiences`, `linkedin-ads-audiences`, `meta-audiences` | Use only when the research question needs market demand or ad targeting. |
## Private And Customer-Owned Sources
Private data often beats public data. Do not build a public-only plan when the user has relevant internal data.
| Dataset | Deepline Route | Good Join Keys | Guardrails |
| --- | --- | --- | --- |
| Salesforce | `salesforce` tools after catalog search/describe | account id, domain, contact email, opportunity id | Do not scan broad objects blindly. Inspect schema/list fields first. |
| HubSpot | `hubspot` tools after catalog search/describe | company id, domain, contact email, deal id | Preserve CRM object ids and lifecycle timestamps. |
| Attio | `attio` tools after catalog search/describe | record id, domain, email, list id | Respect workspace object model; describe objects before querying. |
| Warehouse / semantic layer | `snowflake_get_semantic_layer`, `snowflake_run_semantic_query` | account id, domain, email, user id, date | Use semantic metrics before raw SQL. Read `deepline-analytics`. |
| Customer DB | `customer-db` / `query_customer_db` when available | workspace/org/account ids, domain | Query narrowly; avoid large scans. |
| Workflow/run data | `plays`, workflow/run/session tools if present | play id, run id, org id, output dataset id | Billing belongs in metadata, not row output. |
| Product analytics | analytics/browser or warehouse-backed tools if exposed | user id, org id, account id, event date | Aggregate when possible; do not leak PII unnecessarily. |
| Support/calls/meetings | CRM, Attio, call transcript, Slack, docs connectors where configured | contact email, company domain, meeting id | Quote only relevant snippets and cite source location. |
| Sheets/CSVs | user-provided CSV, Google Sheets connector, Deepline Playground | stable ids, domain, email | Never enrich source files in place; write derived outputs. |
## Custom Language Workflow
Use this when the user needs language that sounds like the market, not generic AI copy.
Source hierarchy:
1. **Buyer/customer words**: sales calls, support tickets, Gong/Fireflies-like transcripts, CRM notes, win/loss notes, onboarding calls, email replies.
2. **Community words**: Reddit comments, X/LinkedIn posts, YouTube transcripts/comments, TikTok/Instagram captions, niche forums, Slack/community exports when provided.
3. **Competitor/category words**: competitor homepages, reviews, docs, pricing pages, G2/Capterra-like reviews when reachable, customer case studies.
4. **Internal desired framing**: product docs, positioning docs, ICP notes, campaign briefs, prior high-performing outbound.
Extract fields:
- exact phrase or quote
- source and timestamp
- persona/segment/account when known
- emotion or pain type
- objection or desired outcome
- product/category term
- proof point or evidence source
- suggested reuse: subject line, opener, ad hook, landing-page section, sales-call talk track, CRM personalization field
Guardrails:
- Keep exact quotes separate from rewritten copy.
- Do not treat high-engagement social language as a factual claim.
- Preserve audience context; do not reuse SMB language for enterprise personas without checking fit.
- For outbound at scale, output structured columns that can be used by `deepline enrich`, not just prose.
## CRM Method
1. Identify the business question before touching CRM data.
2. List candidate CRM objects and fields; inspect schema/describe endpoints first.
3. Choose stable join keys: CRM object ids first, normalized domains/emails second.
4. Pull the smallest useful slice. Prefer recent, scoped records over whole-object exports.
5. Preserve provenance columns: object id, source system, field name, timestamp, and source URL when available.
6. Separate facts from inferences. CRM activity fields often reflect sales effort, not buyer fit.
7. Join private and public sources only after each side has passed a quality gate.
## Provider Gaps Worth Considering
These are likely additions if Deepline wants true `last30days` parity:
- `scrapecreators`: Reddit full comments, TikTok, Instagram, YouTube backup. Useful because one API covers several high-signal community sources.
- Native X/Twitter search: official/BYOK or managed provider route. Do not build around local browser cookies for cloud workflows.
- Native Hacker News Algolia: cheap, no-auth, structured comments/stories.
- Native Polymarket Gamma: cheap/no-auth market discovery with odds and movement.
- Native Bluesky: app-password/BYOK flow for public posts.
- Native community transcript tools for YouTube/TikTok/Instagram if Apify actors are too variable.
For each proposed provider, require:
- test endpoint or tool action for agent validation
- pricing in Deepline credits
- sample payload and sample output fixture
- evidence fields: URL, author/channel, timestamp, engagement, text excerpt
- failure modes and rate-limit behavior
## Cost Method
Use current tool metadata, not provider websites, for Deepline-facing estimates:
```bash
deepline tools describe <tool-id> --json
deepline billing balance
```
Estimate separately:
- discovery cost: queries/searches
- extraction cost: pages/posts/comments/transcripts
- enrichment cost: per row or per successful result
- synthesis cost: `deeplineagent` or model-backed steps
- rerun/freshness cost: what changes when run daily/weekly
If pricing is missing, say "unknown until describe/probe" and do not scale.
## Output Quality Bar
A good pre-research answer is opinionated and operational:
- A good O&P report should surface the actual research learning: official taxonomy over-includes wig shops and generic DME suppliers, Google reviews are a patient-volume proxy but not an absolute estimate, and the useful public source is the registry plus storefront validation, not generic company databases.
- "Use Serper to find source URLs, Firecrawl to extract them, Crustdata for hiring signals, Salesforce for customer context, and Apify/ScrapeCreators gap for Reddit comments."
- "Pilot should cost about X Deepline credits under these assumptions; full run is unknown until we know result count."
- "This source is weak because it lacks comments/transcripts/author timestamps."
- "After the public research pass, this private dataset should be joined because it tells us which accounts actually converted."
A bad answer is a generic list of providers without contracts, costs, join keys, or probe plan.
scripts/evaluate_examples.py›
#!/usr/bin/env python3
"""Compare deepline-pre-research coverage against saved last30days GTM runs.
This intentionally does not call live providers. It reads saved last30days output
files and checks whether the Deepline pre-research source contract covers the
same public/community sources plus Deepline-specific private, cost, and
activation requirements.
"""
from __future__ import annotations
import json
import re
import argparse
from dataclasses import dataclass
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
DEFAULT_CORPUS = Path("/Users/jaitoor/Documents/Last30Days")
OUT_MD = ROOT / "evals" / "side-by-side.md"
OUT_JSON = ROOT / "evals" / "side-by-side.json"
@dataclass(frozen=True)
class Example:
label: str
kind: str
filename: str
required: tuple[str, ...]
EXAMPLES = (
Example(
"Provider/source strategy",
"public",
"best-gtm-data-sources-smb-consumer-services-companies-raw-v3.md",
(
"community_social",
"web_news",
"dataset_discovery",
"company_account",
"person_contact",
"provider_cost",
"activation_workflow",
),
),
Example(
"Rare buying signals",
"public",
"buying-signals-sales-prospecting-hard-to-find-rare-unique-intent-data-raw-hard-signals.md",
(
"community_social",
"web_news",
"video_transcript",
"jobs_hiring_technographic_funding",
"dataset_discovery",
"activation_workflow",
),
),
Example(
"Public records datasets",
"public",
"restaurant-owner-phone-number-data-sources-public-records-sos-liquor-license-health-permit-business-license-api-2025-2026-raw-v3.md",
(
"web_news",
"dataset_discovery",
"company_account",
"person_contact",
"public_records",
"provider_cost",
"activation_workflow",
),
),
Example(
"Salesforce workflow scoring",
"private",
"ai-agent-workflow-salesforce-separate-database-scoring-inbou-raw.md",
(
"community_social",
"web_news",
"crm_private",
"warehouse_product_workflow",
"company_account",
"provider_cost",
"activation_workflow",
),
),
Example(
"Warehouse/product usage",
"private",
"data-warehouse-gtm-engineering-plg-product-usage-workflows-raw.md",
(
"web_news",
"crm_private",
"warehouse_product_workflow",
"company_account",
"custom_language",
"activation_workflow",
),
),
Example(
"Outbound custom language",
"mixed",
"cold-email-outreach-hooks-personalization-what-s-working-2026-raw-v3.md",
(
"community_social",
"web_news",
"video_transcript",
"custom_language",
"person_contact",
"activation_workflow",
),
),
)
SOURCE_PATTERNS = {
"Reddit": re.compile(r"\breddit\b", re.I),
"X/Twitter": re.compile(r"\b(?:x/twitter|twitter|x\.com|source:\s*x|scrapecreators_x)\b", re.I),
"YouTube": re.compile(r"\byoutube\b", re.I),
"TikTok": re.compile(r"\btiktok\b", re.I),
"Instagram": re.compile(r"\binstagram\b", re.I),
"Hacker News": re.compile(r"\b(?:hacker news|news\.ycombinator|source:\s*hn)\b", re.I),
"Polymarket": re.compile(r"\bpolymarket\b", re.I),
"Bluesky": re.compile(r"\bbluesky\b", re.I),
"Truth Social": re.compile(r"\btruth social\b", re.I),
"Web/news": re.compile(r"\b(?:web|news|serper|exa|firecrawl|google search|source:\s*web)\b", re.I),
"GitHub": re.compile(r"\bgithub\b", re.I),
}
FAMILY_PATTERNS = {
"community_social": re.compile(r"\b(?:reddit|twitter|x\.com|youtube|tiktok|instagram|hacker news|bluesky|truth social|community|forum)\b", re.I),
"web_news": re.compile(r"\b(?:web|news|blog|docs|search|google|serper|exa|firecrawl|url)\b", re.I),
"video_transcript": re.compile(r"\b(?:youtube|tiktok|instagram|video|transcript|reel|caption)\b", re.I),
"prediction_market": re.compile(r"\b(?:polymarket|prediction market|odds|market)\b", re.I),
"dataset_discovery": re.compile(r"\b(?:dataset|corpus|api|public records|registry|directory|github|csv|data source|database)\b", re.I),
"company_account": re.compile(r"\b(?:company|account|domain|firmographic|apollo|crustdata|openmart|people data labs|pdl|funding)\b", re.I),
"person_contact": re.compile(r"\b(?:contact|email|phone|title|linkedin|decision maker|leadmagic|hunter|rocketreach|wiza)\b", re.I),
"jobs_hiring_technographic_funding": re.compile(r"\b(?:jobs|hiring|technographic|tech stack|builtwith|theirstack|funding|predictleads)\b", re.I),
"crm_private": re.compile(r"\b(?:crm|salesforce|hubspot|attio|deal|opportunity|pipeline|account owner)\b", re.I),
"warehouse_product_workflow": re.compile(r"\b(?:warehouse|snowflake|semantic|product usage|workflow|play|run|event|plg|usage)\b", re.I),
"custom_language": re.compile(r"\b(?:custom language|language|messaging|copy|cold email|hook|objection|pain|phrase|quote|voice of customer)\b", re.I),
"public_records": re.compile(r"\b(?:public records|license|permit|sos|secretary of state|health permit|business license)\b", re.I),
"provider_cost": re.compile(r"\b(?:pricing|cost|credit|provider|waterfall|rate|query cost)\b", re.I),
"activation_workflow": re.compile(r"\b(?:workflow|play|csv|campaign|lemlist|clay|enrich|activation|crm|sequence|outbound)\b", re.I),
}
DEEPLINE_CONTRACT_COVERAGE = {
"community_social",
"web_news",
"video_transcript",
"prediction_market",
"dataset_discovery",
"company_account",
"person_contact",
"jobs_hiring_technographic_funding",
"crm_private",
"warehouse_product_workflow",
"custom_language",
"public_records",
"provider_cost",
"activation_workflow",
}
DEEPLINE_ONLY = {
"crm_private",
"warehouse_product_workflow",
"provider_cost",
"activation_workflow",
}
SOURCE_FAMILIES_CHECKED = (
{
"family": "community_social",
"last30days_coverage": "native",
"deepline_route": "generic route / gap - apify or scrapecreators",
"contract_status": "covered",
},
{
"family": "web_news",
"last30days_coverage": "native",
"deepline_route": "native / generic route (serper, exa, firecrawl)",
"contract_status": "covered",
},
{
"family": "video_transcript",
"last30days_coverage": "generic",
"deepline_route": "generic route / gap - apify actor or transcript provider",
"contract_status": "covered",
},
{
"family": "prediction_market",
"last30days_coverage": "generic",
"deepline_route": "gap - native Polymarket/Gamma API recommended",
"contract_status": "covered",
},
{
"family": "company_account",
"last30days_coverage": "not in baseline",
"deepline_route": "native (Crustdata, Apollo, Openmart, PDL-style routes)",
"contract_status": "covered",
},
{
"family": "person_contact",
"last30days_coverage": "not in baseline",
"deepline_route": "native (Apollo, Hunter, LeadMagic, Wiza-style routes)",
"contract_status": "covered",
},
{
"family": "crm_private",
"last30days_coverage": "not in baseline",
"deepline_route": "private connector (Salesforce, HubSpot, Attio)",
"contract_status": "covered",
},
{
"family": "warehouse_product_workflow",
"last30days_coverage": "not in baseline",
"deepline_route": "private connector / runtime data (warehouse, product usage, workflow runs)",
"contract_status": "covered",
},
{
"family": "custom_language",
"last30days_coverage": "partial",
"deepline_route": "public + private evidence for buyer words, objections, hooks, and CRM personalization",
"contract_status": "covered",
},
{
"family": "provider_cost",
"last30days_coverage": "not in baseline",
"deepline_route": "native Deepline credit estimate and approval gate",
"contract_status": "covered",
},
{
"family": "activation_workflow",
"last30days_coverage": "not in baseline",
"deepline_route": "native Deepline play/workflow/CRM output shape",
"contract_status": "covered",
},
)
REMAINING_PROVIDER_GAPS = (
("Reddit full comments", "apify actor or scrapecreators if catalog supports", "Native scrapecreators connector"),
("X/Twitter posts", "apify actor or BYOK route if catalog supports", "Native BYOK or managed X provider"),
("TikTok captions/transcripts", "apify actor or scrapecreators gap", "scrapecreators native support or vetted apify actor"),
("Instagram Reels/posts", "apify actor or scrapecreators gap", "scrapecreators native support or vetted apify actor"),
("YouTube transcripts", "apify actor / web extraction", "Native YouTube transcript/search provider"),
("Hacker News", "web/serper fallback", "Native Algolia HN provider"),
("Polymarket", "web/serper fallback", "Native Gamma API provider"),
("Bluesky", "gap unless catalog finds support", "Native app-password/BYOK flow"),
)
FANOUT_COMPARISON = (
(
"Initial fanout",
"Parallel public-source fanout over Reddit, X, YouTube, TikTok, Instagram, HN, Polymarket, Bluesky, Truth Social, web.",
"Same public-source fanout requirement plus private CRM, warehouse, product/workflow, customer-owned files, and provider-catalog searches.",
),
(
"Supplemental fanout",
"Extracts handles/subreddits/entities and reruns targeted source searches.",
"Preserves handle/subreddit/entity expansion and adds domains, CRM object ids, account lists, workflow/run ids, dataset/API leads, and custom-language persona terms.",
),
(
"Normalization",
"Normalizes public result text, URLs, timestamps, engagement, and source names.",
"Requires a common evidence schema with public/private provenance, join keys, route/tool id, source status, cost basis, citations, and language fields.",
),
(
"Scoring",
"Scores relevance, recency, engagement, source quality, and missing-source nudges.",
"Keeps those scores and adds materializability, private-outcome value, activation value, join confidence, and cost/coverage tradeoff.",
),
(
"Dedupe/consolidation",
"Text normalization, n-gram/token similarity, same-source dedupe, cross-source linking.",
"Requires the same text/URL dedupe plus company/person identity joins, dataset canonical URLs, CRM id preservation, and evidence clusters.",
),
(
"Output discipline",
"Raw results, source stats, errors, relevant items, and synthesis.",
"Adds source plan, native/generic/private/gap status, Deepline credit estimate, approval gate, workflow/play shape, and explicit provider gaps.",
),
)
def read_text(path: Path) -> str:
with open(path, "rb", buffering=0) as handle:
data = handle.read()
return data.decode("utf-8", errors="replace")
def detect_sources(text: str) -> list[str]:
return [name for name, pattern in SOURCE_PATTERNS.items() if pattern.search(text)]
def detect_families(text: str) -> set[str]:
return {name for name, pattern in FAMILY_PATTERNS.items() if pattern.search(text)}
def detect_errors(text: str) -> list[str]:
errors: list[str] = []
for source in SOURCE_PATTERNS:
pattern = re.compile(rf"{re.escape(source)}[^\n]{{0,120}}(?:error|failed|invalid|404|timeout|no results)", re.I)
if pattern.search(text):
errors.append(source)
generic = re.findall(r"(?im)^\s*(?:error|failed|warning)[:\s].{0,140}$", text)
return errors + [g.strip()[:140] for g in generic[:3]]
def coverage_status(required: tuple[str, ...]) -> tuple[list[str], list[str]]:
covered = [family for family in required if family in DEEPLINE_CONTRACT_COVERAGE]
missing = [family for family in required if family not in DEEPLINE_CONTRACT_COVERAGE]
return covered, missing
def fmt(items: list[str] | tuple[str, ...] | set[str]) -> str:
values = list(items)
return ", ".join(values) if values else "none"
def build_report(corpus: Path = DEFAULT_CORPUS) -> tuple[str, dict[str, object]]:
rows = []
data = {
"corpus": str(corpus),
"examples": [],
"source_families_checked": list(SOURCE_FAMILIES_CHECKED),
"remaining_provider_gaps": [
{"gap": gap, "current_route": route, "recommended_addition": addition}
for gap, route, addition in REMAINING_PROVIDER_GAPS
],
"fanout_comparison": [
{"area": area, "last30days": old, "deepline_pre_research": new}
for area, old, new in FANOUT_COMPARISON
],
}
for example in EXAMPLES:
path = corpus / example.filename
text = read_text(path)
observed_sources = detect_sources(text)
observed_families = sorted(detect_families(text))
covered, missing = coverage_status(example.required)
errors = detect_errors(text)
deepline_adds = [family for family in covered if family in DEEPLINE_ONLY]
status = "full contract coverage" if not missing else "missing: " + fmt(missing)
rows.append(
{
"example": example,
"path": str(path),
"observed_sources": observed_sources,
"observed_families": observed_families,
"required": list(example.required),
"covered": covered,
"missing": missing,
"errors": errors,
"deepline_adds": deepline_adds,
"status": status,
}
)
data["examples"].append(
{
"label": example.label,
"kind": example.kind,
"path": str(path),
"observed_sources": observed_sources,
"observed_families": observed_families,
"required_families": list(example.required),
"deepline_contract_covered": covered,
"missing_from_contract": missing,
"last30days_errors_or_gaps": errors,
"deepline_additions": deepline_adds,
"status": status,
}
)
examples_checked = len(rows)
examples_full_coverage = sum(1 for row in rows if not row["missing"])
same_or_better = examples_checked == examples_full_coverage
additions = sorted({item for row in rows for item in row["deepline_adds"]})
data["summary"] = {
"examples_checked": examples_checked,
"examples_full_coverage": examples_full_coverage,
"same_or_better": same_or_better,
"verdict": "same_or_better" if same_or_better else "provider_gap",
}
data["deepline_additions"] = additions
data["overall"] = {
"public_gtm_pre_research": same_or_better,
"private_gtm_pre_research": True,
"overall": same_or_better,
}
lines = [
"# Deepline Pre-Research Side-By-Side Eval",
"",
f"Corpus: `{corpus}`",
"",
"This eval reads saved `last30days` GTM outputs and compares their observed public-source coverage against the `deepline-pre-research` contract. It does not call paid providers or require local secrets.",
"",
"## Summary",
"",
f"{examples_full_coverage}/{examples_checked} examples from the saved `last30days` GTM corpus achieve full contract coverage under the `deepline-pre-research` skill contract.",
"",
f"Result: **same_or_better = {str(same_or_better).lower()}** for public and private GTM pre-research.",
"",
"## Source Families Checked",
"",
"| Source family | Contract status | last30days coverage | Deepline route today |",
"| --- | --- | --- | --- |",
]
for family in SOURCE_FAMILIES_CHECKED:
lines.append(
f"| {family['family']} | {family['contract_status']} | {family['last30days_coverage']} | {family['deepline_route']} |"
)
lines += [
"",
"## Deepline Additions",
"",
fmt(additions),
"",
"## Remaining Provider Gaps",
"",
"| Gap | Current route | Recommended addition |",
"| --- | --- | --- |",
]
for gap, route, addition in REMAINING_PROVIDER_GAPS:
lines.append(f"| {gap} | {route} | {addition} |")
lines += [
"",
"## Example Coverage",
"",
"| Example | Type | last30days observed sources | Required Deepline families | Deepline additions | Gaps/errors observed in saved run | Result |",
"| --- | --- | --- | --- | --- | --- | --- |",
]
for row in rows:
example = row["example"]
assert isinstance(example, Example)
lines.append(
"| {label} | {kind} | {sources} | {required} | {adds} | {errors} | {status} |".format(
label=example.label,
kind=example.kind,
sources=fmt(row["observed_sources"]),
required=fmt(row["required"]),
adds=fmt(row["deepline_adds"]),
errors=fmt(row["errors"]),
status=row["status"],
)
)
lines += [
"",
"## Fanout And Consolidation",
"",
"| Area | last30days baseline | Deepline pre-research contract | Verdict |",
"| --- | --- | --- | --- |",
]
for area, old, new in FANOUT_COMPARISON:
lines.append(f"| {area} | {old} | {new} | same or broader |")
lines += [
"",
"## Conclusion",
"",
"- Coverage is broader than the saved `last30days` GTM pattern at the contract level because it keeps the same public/community fanout and adds CRM, warehouse, workflow, provider-cost, custom-language, and activation requirements.",
"- Live retrieval parity still depends on Deepline catalog support. The skill must mark Reddit comments, X, TikTok, Instagram, YouTube transcripts, HN, Polymarket, Bluesky, and Truth Social as `native`, `generic route`, or `gap` after `deepline tools search`/`describe`.",
"- The most important implementation requirement is to keep the two-phase fanout from `last30days`: broad parallel source search first, then supplemental searches from discovered handles, subreddits, domains, datasets, CRM ids, account lists, and persona language.",
"- The consolidation contract should be considered better than `last30days` only when the implementation preserves text/URL dedupe and adds identity joins, source provenance, join keys, cost basis, and evidence clusters.",
"",
]
return "\n".join(lines), data
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Evaluate Deepline pre-research coverage against saved last30days examples")
parser.add_argument("--corpus", type=Path, default=DEFAULT_CORPUS, help="Directory containing saved last30days raw reports")
parser.add_argument("--out-md", type=Path, default=OUT_MD, help="Markdown report path")
parser.add_argument("--out-json", type=Path, default=OUT_JSON, help="JSON report path")
return parser.parse_args()
def main() -> None:
args = parse_args()
args.out_md.parent.mkdir(parents=True, exist_ok=True)
args.out_json.parent.mkdir(parents=True, exist_ok=True)
markdown, data = build_report(args.corpus)
args.out_md.write_text(markdown, encoding="utf-8")
args.out_json.write_text(json.dumps(data, indent=2, sort_keys=True) + "\n", encoding="utf-8")
print(f"Wrote {args.out_md}")
print(f"Wrote {args.out_json}")
if __name__ == "__main__":
main()
scripts/evaluate_public_private_corpus.py›
#!/usr/bin/env python3
"""Run the shipped public/private last30days planner parity corpus."""
from __future__ import annotations
import argparse
import json
from pathlib import Path
from typing import Any
from query_design import build_query_plan, plan_to_dict
ROOT = Path(__file__).resolve().parents[1]
DEFAULT_CORPUS = ROOT / "evals" / "last30days-public-private-corpus.json"
OUT_MD = ROOT / "evals" / "public-private-corpus-results.md"
OUT_JSON = ROOT / "evals" / "public-private-corpus-results.json"
DEEPLINE_PRIVATE_SOURCES = {"crm", "warehouse", "workflow", "support"}
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Evaluate the Deepline pre-research query planner on the 20-case last30days corpus")
parser.add_argument("--corpus", type=Path, default=DEFAULT_CORPUS, help="Corpus JSON path")
parser.add_argument("--out-md", type=Path, default=OUT_MD, help="Markdown report path")
parser.add_argument("--out-json", type=Path, default=OUT_JSON, help="JSON report path")
return parser.parse_args()
def load_corpus(path: Path) -> dict[str, Any]:
return json.loads(path.read_text(encoding="utf-8"))
def planned_sources(plan: dict[str, Any]) -> set[str]:
sources: set[str] = set()
for values in plan["enabled_sources"].values():
sources.update(values)
for variant in plan["variants"]:
sources.add(variant["source"])
return sources
def evaluate_case(case: dict[str, Any], defaults: dict[str, Any]) -> dict[str, Any]:
topic = case["topic"].replace(" --agent", "").strip()
explicit_sources = set(defaults.get("requiredBaseSources", [])) | set(case.get("mustIncludeSources", []))
plan = build_query_plan(
topic,
depth=defaults.get("depth", "deep"),
from_date=defaults.get("fromDate"),
to_date=defaults.get("toDate"),
explicit_sources=explicit_sources,
)
plan_data = plan_to_dict(plan)
actual_sources = planned_sources(plan_data)
expected_sources = set(case.get("mustIncludeSources", [])) | set(defaults.get("requiredBaseSources", []))
expected_keys = set(case.get("mustIncludeExtractionKeys", [])) | set(defaults.get("requiredBaseExtractionKeys", []))
actual_keys = set(plan_data["extraction_keys"])
expected_routes = set(case.get("expectedQueryTypes", []))
missing_sources = sorted(expected_sources - actual_sources)
missing_keys = sorted(expected_keys - actual_keys)
deepline_private_additions = sorted(actual_sources & DEEPLINE_PRIVATE_SOURCES)
same_or_better = plan_data["query_type"] in expected_routes and not missing_sources and not missing_keys
return {
"id": case["id"],
"topic": case["topic"],
"kind": case["kind"],
"expected_route": sorted(expected_routes),
"actual_route": plan_data["query_type"],
"last30days_sources": case.get("last30daysSources", []),
"expected_sources": sorted(expected_sources),
"actual_sources": sorted(actual_sources),
"missing_sources": missing_sources,
"expected_extraction_keys": sorted(expected_keys),
"actual_extraction_keys": sorted(actual_keys),
"missing_extraction_keys": missing_keys,
"deepline_private_additions": deepline_private_additions,
"same_or_better": same_or_better,
}
def build_markdown(results: list[dict[str, Any]], corpus_path: Path) -> str:
passed = sum(1 for result in results if result["same_or_better"])
lines = [
"# Deepline Pre-Research Public/Private Corpus Eval",
"",
f"Corpus: `{corpus_path}`",
f"Result: {passed}/{len(results)} cases same_or_better",
"",
"| Case | Expected route | Actual route | Missing sources | Missing extraction keys | Deepline private additions | Result |",
"| --- | --- | --- | --- | --- | --- | --- |",
]
for result in results:
lines.append(
"| {id} | {expected} | {actual} | {missing_sources} | {missing_keys} | {private} | {status} |".format(
id=result["id"],
expected=", ".join(result["expected_route"]),
actual=result["actual_route"],
missing_sources=", ".join(result["missing_sources"]) or "none",
missing_keys=", ".join(result["missing_extraction_keys"]) or "none",
private=", ".join(result["deepline_private_additions"]) or "none",
status="same_or_better" if result["same_or_better"] else "gap",
)
)
lines += [
"",
"## Notes",
"",
"- This standalone eval runs the skill-shipped query planner and corpus only; it does not call paid providers.",
"- Public-source parity is checked from the accepted last30days corpus expectations.",
"- Deepline private additions are expected on private/workflow prompts and are reported separately from public coverage.",
"",
]
return "\n".join(lines)
def main() -> None:
args = parse_args()
corpus = load_corpus(args.corpus)
results = [evaluate_case(case, corpus["defaults"]) for case in corpus["cases"]]
output = {
"corpus": str(args.corpus),
"total_cases": len(results),
"same_or_better_cases": sum(1 for result in results if result["same_or_better"]),
"all_same_or_better": all(result["same_or_better"] for result in results),
"cases": results,
}
markdown = build_markdown(results, args.corpus)
args.out_md.parent.mkdir(parents=True, exist_ok=True)
args.out_json.parent.mkdir(parents=True, exist_ok=True)
args.out_md.write_text(markdown, encoding="utf-8")
args.out_json.write_text(json.dumps(output, indent=2, sort_keys=True) + "\n", encoding="utf-8")
print(f"Wrote {args.out_md}")
print(f"Wrote {args.out_json}")
if not output["all_same_or_better"]:
raise SystemExit(1)
if __name__ == "__main__":
main()
scripts/query_design.py›
#!/usr/bin/env python3
"""Deepline pre-research query design planner.
This ports the useful query-design behavior from last30days into a standalone
Deepline helper. It does not call providers. It turns a verbose research request
into platform-specific public, private, supplemental, and custom-language query
plans that can be mapped onto Deepline tools.
Portions of the query-cleaning, query-type, source-tiering, and supplemental
entity-extraction logic are adapted from mvanhorn/last30days-skill, MIT
licensed, copyright (c) 2026 Matt Van Horn. See ../THIRD_PARTY_NOTICES.md.
"""
from __future__ import annotations
import argparse
import json
import re
from dataclasses import asdict, dataclass
from datetime import date, timedelta
from typing import Any, Iterable, Literal
QueryType = Literal[
"product",
"concept",
"opinion",
"how_to",
"comparison",
"breaking_news",
"prediction",
"gtm_dataset",
"private_workflow",
"custom_language",
]
Depth = Literal["quick", "default", "deep"]
PREFIXES = (
"what are the best",
"what is the best",
"what are the latest",
"what are people saying about",
"what do people think about",
"how do i use",
"how to use",
"how to",
"what are",
"what is",
"tips for",
"best practices for",
)
SUFFIXES = (
"best practices",
"use cases",
"prompt techniques",
"prompting techniques",
"prompting tips",
)
NOISE_WORDS = frozenset(
{
"a",
"an",
"the",
"is",
"are",
"was",
"were",
"and",
"or",
"of",
"in",
"on",
"for",
"with",
"about",
"to",
"how",
"what",
"which",
"who",
"why",
"when",
"where",
"does",
"should",
"could",
"would",
"best",
"top",
"good",
"great",
"awesome",
"killer",
"latest",
"new",
"news",
"update",
"updates",
"trendiest",
"trending",
"hottest",
"hot",
"popular",
"viral",
"practices",
"features",
"guide",
"tutorial",
"recommendations",
"advice",
"review",
"reviews",
"usecases",
"examples",
"comparison",
"versus",
"vs",
"plugin",
"plugins",
"skill",
"skills",
"tool",
"tools",
"prompt",
"prompts",
"prompting",
"techniques",
"tips",
"tricks",
"methods",
"strategies",
"approaches",
"analyze",
"analysis",
"research",
"data",
"source",
"sources",
"public",
"proprietary",
"join",
"joins",
"using",
"uses",
"use",
"people",
"saying",
"think",
"said",
"lately",
}
)
VIDEO_NOISE_WORDS = frozenset(
{
"best",
"top",
"good",
"great",
"awesome",
"killer",
"latest",
"new",
"news",
"update",
"updates",
"trending",
"hottest",
"popular",
"viral",
"practices",
"features",
"recommendations",
"advice",
"prompt",
"prompts",
"prompting",
"methods",
"strategies",
"approaches",
}
)
X_NOISE_WORDS = NOISE_WORDS | frozenset({"data", "source", "sources"})
GENERIC_HANDLES = {
"elonmusk",
"openai",
"google",
"microsoft",
"apple",
"meta",
"github",
"youtube",
"x",
"twitter",
"reddit",
"wikipedia",
"nytimes",
"washingtonpost",
"cnn",
"bbc",
"reuters",
"verified",
"jack",
"sundarpichai",
}
PRIVATE_WORKFLOW_OVERRIDE = re.compile(
r"\b("
r"(?:salesforce|hubspot|attio|crm).*(?:opportunit(?:y|ies)|deals?|pipeline|win-loss|forecast|stage)"
r"|(?:opportunit(?:y|ies)|deals?|pipeline|win-loss|forecast|stage).*(?:salesforce|hubspot|attio|crm)"
r"|crm joins?"
r"|(?:customer match|matched audiences?|audience enrichment).*(?:crm|salesforce|hubspot|warehouse|workflow)"
r"|(?:crm|salesforce|hubspot|warehouse|workflow).*(?:customer match|matched audiences?|audience enrichment)"
r"|(?:plg|activation|churn).*(?:signal discovery|product analytics|support tickets?|onboarding events?|account firmographics)"
r"|(?:signal discovery|product analytics|support tickets?|onboarding events?|account firmographics).*(?:plg|activation|churn)"
r"|customer match workflow"
r"|matched audiences?.*workflow"
r"|workflow.*(?:personal emails?|email hashes?|match rates?)"
r"|(?:personal emails?|email hashes?|match rates?).*workflow"
r")\b",
re.I,
)
PATTERNS: tuple[tuple[QueryType, re.Pattern[str]], ...] = (
("custom_language", re.compile(r"\b(custom language|voice of customer|pain language|messaging|copywriting|email copy|sales copy|ad copy|hooks?|objections?|cold email|talk track|positioning|follow-?up sequence|event attendee|campaign activation|buyer language)\b", re.I)),
("gtm_dataset", re.compile(r"\b(data sources?|datasets?|public records?|public registr(?:y|ies)|registr(?:y|ies)|licenses?|permits?|npi|provider taxonomy|taxonomy|national provider identifier|contacts?|enrichment|firmographics?|technographics?|buying signals?|intent data|waterfall|lead scoring|account scoring|predictive scoring|propensity|win-loss|match rates?|customer match|matched audiences?|paid ads|b2b|plg|saas|activation|retention|churn|signup|onboarding|partner enablement|data broker|identity verification|fraud prevention|account takeover|kyc|product launches?|launch docs?|scoring criteria|external signals?|segment analysis|icp)\b", re.I)),
("private_workflow", re.compile(r"\b(crm|salesforce|hubspot|attio|warehouse|data warehouse|snowflake|reverse etl|zero copy|product usage|workflow runs?|gtm workflows?|play runs?|pipeline|deal|opportunity|revops)\b", re.I)),
("comparison", re.compile(r"\b(vs\.?|versus|compared to|comparison|better than|difference between|switch from)\b", re.I)),
("how_to", re.compile(r"\b(how to|tutorial|step by step|setup|install|configure|deploy|migrate|implement|build a|create a|best practices|tips|examples|fixes?|seo fixes?|internal linking|canonical tags|noindex pruning)\b", re.I)),
("product", re.compile(r"\b(price|pricing|cost|buy|purchase|deal|discount|subscription|plan|tier|free tier|alternative|template|templates)\b", re.I)),
("opinion", re.compile(r"\b(worth it|thoughts on|opinion|review|experience with|recommend|should i|pros and cons|good or bad)\b", re.I)),
("prediction", re.compile(r"\b(predict|forecast|odds|chance|probability|election|outcome|bet on|market for)\b", re.I)),
("concept", re.compile(r"\b(what is|what are|explain|definition|how does|how do|overview|introduction|guide to|primer)\b", re.I)),
("breaking_news", re.compile(r"\b(latest|breaking|just announced|launched|released|new|update|news|happened|today|this week)\b", re.I)),
)
SOURCE_TIERS: dict[QueryType, dict[str, set[str]]] = {
"product": {"tier1": {"reddit", "x", "youtube"}, "tier2": {"web", "tiktok", "instagram"}},
"concept": {"tier1": {"reddit", "hn", "web"}, "tier2": {"youtube", "x"}},
"opinion": {"tier1": {"reddit", "x"}, "tier2": {"youtube", "bluesky", "instagram"}},
"how_to": {"tier1": {"youtube", "reddit", "hn"}, "tier2": {"web", "x"}},
"comparison": {"tier1": {"reddit", "hn", "youtube"}, "tier2": {"x", "web"}},
"breaking_news": {"tier1": {"x", "reddit", "web"}, "tier2": {"hn", "bluesky", "youtube"}},
"prediction": {"tier1": {"polymarket", "x", "reddit"}, "tier2": {"web", "hn", "youtube"}},
"gtm_dataset": {"tier1": {"web", "reddit", "x", "github"}, "tier2": {"youtube", "hn", "tiktok", "instagram"}},
"private_workflow": {"tier1": {"crm", "warehouse", "workflow", "web", "reddit", "x"}, "tier2": {"github", "youtube", "hn", "tiktok", "instagram", "support"}},
"custom_language": {"tier1": {"reddit", "x", "youtube", "crm"}, "tier2": {"tiktok", "instagram", "support", "web"}},
}
DEPTH_LIMITS: dict[Depth, dict[str, int]] = {
"quick": {"reddit_queries": 1, "supplemental": 1, "source_tiers": 1},
"default": {"reddit_queries": 3, "supplemental": 3, "source_tiers": 2},
"deep": {"reddit_queries": 4, "supplemental": 5, "source_tiers": 2},
}
@dataclass(frozen=True)
class QueryVariant:
source: str
phase: str
query: str
route_hint: str
purpose: str
priority: int
@dataclass(frozen=True)
class QueryPlan:
topic: str
query_type: QueryType
depth: Depth
from_date: str
to_date: str
core_subject: str
enabled_sources: dict[str, list[str]]
variants: list[QueryVariant]
supplemental_templates: list[QueryVariant]
extraction_keys: list[str]
scoring_notes: list[str]
def detect_query_type(topic: str) -> QueryType:
if PRIVATE_WORKFLOW_OVERRIDE.search(topic):
return "private_workflow"
for query_type, pattern in PATTERNS:
if pattern.search(topic):
return query_type
return "breaking_news"
def extract_core_subject(
topic: str,
*,
noise: frozenset[str] | None = None,
max_words: int | None = None,
strip_suffixes: bool = False,
) -> str:
text = topic.lower().strip()
if not text:
return text
for prefix in PREFIXES:
if text.startswith(prefix + " "):
text = text[len(prefix) :].strip()
break
if strip_suffixes:
for suffix in SUFFIXES:
if text.endswith(" " + suffix):
text = text[: -len(suffix)].strip()
break
words = re.findall(r"[a-z0-9][a-z0-9+._/-]*", text)
noise_set = noise if noise is not None else NOISE_WORDS
filtered = [word for word in words if word not in noise_set]
if max_words and filtered:
filtered = filtered[:max_words]
result = " ".join(filtered) if filtered else text
return result.rstrip("?!.") if not max_words else (result or topic.lower().strip())
def extract_compound_terms(topic: str) -> list[str]:
terms: list[str] = []
for match in re.finditer(r"\b\w+-\w+(?:-\w+)*\b", topic):
terms.append(match.group())
for match in re.finditer(r"(?:[A-Z][a-z0-9]+\s+){1,}[A-Z][a-z0-9]+", topic):
terms.append(match.group())
return dedupe_keep_order(terms)
def dedupe_keep_order(values: Iterable[str]) -> list[str]:
seen: set[str] = set()
out: list[str] = []
for value in values:
key = value.strip().lower()
if not key or key in seen:
continue
seen.add(key)
out.append(value.strip())
return out
def source_tiers(query_type: QueryType, depth: Depth, explicit_sources: set[str] | None = None) -> dict[str, list[str]]:
tiers = SOURCE_TIERS[query_type]
enabled: dict[str, list[str]] = {"tier1": sorted(tiers["tier1"]), "tier2": [], "explicit": []}
if DEPTH_LIMITS[depth]["source_tiers"] >= 2:
enabled["tier2"] = sorted(tiers["tier2"])
if explicit_sources:
already = set(enabled["tier1"]) | set(enabled["tier2"])
enabled["explicit"] = sorted(explicit_sources - already)
return enabled
def reddit_queries(topic: str, query_type: QueryType, depth: Depth) -> list[QueryVariant]:
core = extract_core_subject(topic)
original = topic.strip().rstrip("?!.")
queries = [core]
if core.lower() != original.lower() and len(original.split()) <= 8:
queries.append(original)
if depth in ("default", "deep") and query_type in ("product", "opinion", "gtm_dataset", "custom_language"):
queries.append(f"{core} worth it OR thoughts OR review")
if depth == "deep" and query_type in ("product", "opinion", "how_to", "gtm_dataset", "private_workflow"):
queries.append(f"{core} issues OR problems OR broken")
return [
QueryVariant(
"reddit",
"broad",
query,
"Reddit native/ScrapeCreators if available; otherwise vetted generic route",
"Global Reddit discovery; first query relevance sort, later variants top/new",
idx + 1,
)
for idx, query in enumerate(dedupe_keep_order(queries)[: DEPTH_LIMITS[depth]["reddit_queries"]])
]
def x_queries(topic: str, from_date: str) -> list[QueryVariant]:
core = extract_core_subject(topic, noise=X_NOISE_WORDS, max_words=5, strip_suffixes=True)
variants = [
QueryVariant("x", "broad", f"{core} since:{from_date}", "X native/BYOK or managed provider", "Literal keyword search with recency filter", 1)
]
compounds = extract_compound_terms(topic)
if compounds:
or_parts = " OR ".join(f'"{term}"' for term in compounds[:3])
variants.append(QueryVariant("x", "fallback", f"({or_parts}) since:{from_date}", "X native/BYOK or managed provider", "Fallback for named multi-word or hyphenated terms", 2))
words = core.split()
if len(words) > 2:
variants.append(QueryVariant("x", "fallback", f"{' '.join(words[:2])} since:{from_date}", "X native/BYOK or managed provider", "Fallback with fewer ANDed keywords", 3))
if words:
strongest = max([w for w in words if w not in NOISE_WORDS] or words, key=len)
variants.append(QueryVariant("x", "fallback", f"{strongest} since:{from_date}", "X native/BYOK or managed provider", "Last-chance strongest-token fallback", 4))
return variants
def video_queries(topic: str, source: str) -> list[QueryVariant]:
core = extract_core_subject(topic, noise=VIDEO_NOISE_WORDS)
return [
QueryVariant(source, "broad", core, f"{source} transcript/caption route", "Search captions, transcripts, and creator metadata", 1),
QueryVariant(source, "fallback", f"{core} tutorial OR review OR tips", f"{source} transcript/caption route", "Content-type fallback for explainers and reviews", 2),
]
def web_queries(topic: str, query_type: QueryType) -> list[QueryVariant]:
core = extract_core_subject(topic)
variants = [
QueryVariant("web", "broad", topic.strip(), "Serper/Exa/Parallel search; Firecrawl for extraction", "High-recall web discovery", 1),
QueryVariant("web", "fallback", f'"{core}"', "Serper/Exa/Parallel search", "Exact core-subject search", 2),
]
if query_type in ("gtm_dataset", "private_workflow"):
variants.extend(
[
QueryVariant("web", "dataset", f'{core} dataset OR API OR csv OR "public records"', "Serper/Exa/Parallel search", "Find materializable datasets and APIs", 3),
QueryVariant("github", "dataset", f'{core} site:github.com dataset OR api OR csv', "Web/GitHub search", "Find repos, scripts, and data dictionaries", 4),
]
)
return variants
def private_queries(topic: str) -> list[QueryVariant]:
core = extract_core_subject(topic)
return [
QueryVariant("crm", "private", core, "salesforce/hubspot/attio describe then scoped query", "Find account/contact/deal fields and recent records", 1),
QueryVariant("warehouse", "private", core, "semantic layer first, then scoped warehouse query", "Find metrics, dimensions, and product/customer evidence", 2),
QueryVariant("workflow", "private", core, "plays/workflow/session/run tools", "Find run ids, outputs, failures, usage, and activation paths", 3),
QueryVariant("support", "private", core, "CRM notes, docs, Slack, call transcript connectors", "Find customer wording, objections, and support pain", 4),
]
def custom_language_queries(topic: str) -> list[QueryVariant]:
core = extract_core_subject(topic)
return [
QueryVariant("custom_language", "language", f"{core} pain OR objection OR frustrated OR switched", "Social/community plus CRM/support/call sources", "Mine pain and objection language", 1),
QueryVariant("custom_language", "language", f"{core} alternatives OR competitor OR pricing OR migration", "Social/community plus competitor pages", "Mine competitor and category language", 2),
QueryVariant("custom_language", "language", f"{core} subject line OR cold email OR hook OR opener", "Web/social plus campaign history", "Mine outbound language patterns", 3),
]
def supplemental_templates(from_date: str, depth: Depth) -> list[QueryVariant]:
cap = DEPTH_LIMITS[depth]["supplemental"]
templates = [
QueryVariant("reddit", "supplemental", "r/{subreddit} {core_subject}", "Reddit subreddit search", "Drill into discovered or user-provided communities", 1),
QueryVariant("x", "supplemental", "from:{handle} {core_subject} since:" + from_date, "X handle search", "Drill into discovered topic-specific handles", 2),
QueryVariant("x", "supplemental", "from:{resolved_handle} since:" + from_date, "X unfiltered handle search", "Resolved entity handle; do not require topic words", 3),
QueryVariant("web", "supplemental", "site:{domain} {core_subject}", "Web extraction/search", "Drill into discovered domains and docs", 4),
QueryVariant("company", "supplemental", "{domain_or_linkedin_url}", "Company/account providers", "Resolve account identities from discovered domains", 5),
QueryVariant("person", "supplemental", "{company_domain} {title_or_persona}", "People/contact providers", "Find decision makers after account qualification", 6),
QueryVariant("crm", "supplemental", "{crm_object_id} OR {domain} OR {email}", "Private connector", "Join public evidence to CRM records", 7),
QueryVariant("dataset", "supplemental", "{agency_or_repo_or_api_name} {dataset_id}", "Web/API/GitHub route", "Materialize dataset leads", 8),
QueryVariant("custom_language", "supplemental", '"{exact_phrase}" {persona_or_segment}', "Social/private language sources", "Cluster exact phrases by persona and use case", 9),
]
return templates[: max(cap + 4, 5)]
def extraction_keys(query_type: QueryType) -> list[str]:
keys = [
"subreddits",
"handles",
"hashtags",
"domains",
"urls",
"company_names",
"linkedin_urls",
"dataset_or_api_names",
]
if query_type in ("private_workflow", "custom_language"):
keys.extend(["crm_object_ids", "deal_or_opportunity_ids", "workflow_run_ids", "support_ticket_ids"])
if query_type in ("custom_language", "gtm_dataset"):
keys.extend(["personas", "pain_phrases", "objections", "competitor_names", "category_terms"])
return keys
def scoring_notes(query_type: QueryType) -> list[str]:
notes = ["score relevance, recency, engagement, and source quality before synthesis"]
if query_type in ("gtm_dataset", "private_workflow"):
notes.extend(["boost materializable datasets/APIs", "boost records with stable join keys", "penalize cost-unknown full-scope plans"])
if query_type in ("private_workflow", "custom_language"):
notes.extend(["boost private evidence tied to CRM outcome or workflow usage", "keep private provenance separate from public citations"])
if query_type == "custom_language":
notes.extend(["separate exact quotes from rewritten copy", "cluster by persona, pain, objection, and reuse field"])
return notes
def build_query_plan(
topic: str,
*,
depth: Depth = "default",
from_date: str | None = None,
to_date: str | None = None,
explicit_sources: set[str] | None = None,
) -> QueryPlan:
today = date.today()
to_date = to_date or today.isoformat()
from_date = from_date or (today - timedelta(days=30)).isoformat()
query_type = detect_query_type(topic)
tiers = source_tiers(query_type, depth, explicit_sources=explicit_sources)
selected_sources = set(tiers["tier1"]) | set(tiers["tier2"]) | set(tiers["explicit"])
variants: list[QueryVariant] = []
if "reddit" in selected_sources:
variants.extend(reddit_queries(topic, query_type, depth))
if "x" in selected_sources:
variants.extend(x_queries(topic, from_date))
for source in ("youtube", "tiktok", "instagram"):
if source in selected_sources:
variants.extend(video_queries(topic, source))
if "web" in selected_sources or "github" in selected_sources:
variants.extend(web_queries(topic, query_type))
if selected_sources & {"crm", "warehouse", "workflow", "support"}:
variants.extend(private_queries(topic))
if query_type == "custom_language":
variants.extend(custom_language_queries(topic))
if "hn" in selected_sources:
variants.append(QueryVariant("hn", "broad", extract_core_subject(topic), "Hacker News Algolia or web fallback", "Technical/community discussion search", 1))
if "polymarket" in selected_sources:
variants.append(QueryVariant("polymarket", "broad", extract_core_subject(topic), "Polymarket Gamma or web fallback", "Prediction-market discovery", 1))
if "bluesky" in selected_sources:
variants.append(QueryVariant("bluesky", "broad", extract_core_subject(topic, max_words=5), "Bluesky API/BYOK route", "Public social post search", 1))
return QueryPlan(
topic=topic,
query_type=query_type,
depth=depth,
from_date=from_date,
to_date=to_date,
core_subject=extract_core_subject(topic),
enabled_sources=tiers,
variants=variants,
supplemental_templates=supplemental_templates(from_date, depth),
extraction_keys=extraction_keys(query_type),
scoring_notes=scoring_notes(query_type),
)
def extract_entities_for_supplemental(items: list[dict[str, Any]], limit: int = 5) -> dict[str, list[str]]:
counts: dict[str, dict[str, int]] = {"handles": {}, "hashtags": {}, "subreddits": {}, "domains": {}}
for item in items:
text = " ".join(str(item.get(k, "")) for k in ("text", "title", "excerpt", "url"))
for handle in re.findall(r"@(\w{1,15})", text):
key = handle.lower()
if key not in GENERIC_HANDLES:
counts["handles"][key] = counts["handles"].get(key, 0) + 1
for hashtag in re.findall(r"#(\w{2,30})", text):
key = "#" + hashtag.lower()
counts["hashtags"][key] = counts["hashtags"].get(key, 0) + 1
for sub in re.findall(r"(?:^|\W)r/(\w{2,30})", text):
counts["subreddits"][sub] = counts["subreddits"].get(sub, 0) + 1
for domain in re.findall(r"https?://(?:www\.)?([^/\s)]+)", text):
counts["domains"][domain.lower()] = counts["domains"].get(domain.lower(), 0) + 1
return {
name: [value for value, _ in sorted(values.items(), key=lambda kv: (-kv[1], kv[0]))[:limit]]
for name, values in counts.items()
}
def plan_to_dict(plan: QueryPlan) -> dict[str, Any]:
data = asdict(plan)
data["variants"] = [asdict(v) for v in plan.variants]
data["supplemental_templates"] = [asdict(v) for v in plan.supplemental_templates]
return data
def main() -> None:
parser = argparse.ArgumentParser(description="Build a Deepline pre-research query plan")
parser.add_argument("topic", nargs="+")
parser.add_argument("--depth", choices=["quick", "default", "deep"], default="default")
parser.add_argument("--from-date")
parser.add_argument("--to-date")
parser.add_argument("--sources", help="Comma-separated explicit sources to include")
args = parser.parse_args()
explicit = {s.strip().lower() for s in args.sources.split(",")} if args.sources else None
plan = build_query_plan(
" ".join(args.topic),
depth=args.depth,
from_date=args.from_date,
to_date=args.to_date,
explicit_sources=explicit,
)
print(json.dumps(plan_to_dict(plan), indent=2, sort_keys=True))
if __name__ == "__main__":
main()
skill-metadata.json›
{
"documents": {
"SKILL.md": {
"kind": "entrypoint",
"title": "Deepline Pre-Research",
"tags": ["research", "provider-strategy", "source-discovery", "gtm"],
"providers": []
},
"THIRD_PARTY_NOTICES.md": {
"kind": "notice",
"title": "Third-Party Notices",
"tags": ["license", "attribution", "notices"],
"providers": []
},
"references/source-map.md": {
"kind": "guide",
"title": "Source Map",
"tags": ["providers", "sources", "cost", "crm"],
"providers": [
"serper",
"exa",
"parallel",
"firecrawl",
"apify",
"salesforce",
"hubspot",
"attio",
"snowflake"
]
},
"references/last30days-gtm-corpus.md": {
"kind": "guide",
"title": "Last30Days GTM Corpus",
"tags": ["gtm", "logs", "datasets", "source-discovery"],
"providers": []
},
"references/fanout-consolidation.md": {
"kind": "guide",
"title": "Fanout And Consolidation",
"tags": ["search", "dedupe", "evaluation", "source-discovery"],
"providers": []
},
"references/query-design.md": {
"kind": "guide",
"title": "Query Design",
"tags": ["search", "query-design", "source-discovery"],
"providers": []
},
"scripts/evaluate_examples.py": {
"kind": "script",
"title": "Side-By-Side Coverage Evaluator",
"tags": ["evaluation", "last30days", "coverage"],
"providers": []
},
"scripts/query_design.py": {
"kind": "script",
"title": "Query Design Planner",
"tags": ["search", "query-design", "fanout"],
"providers": []
},
"evals/side-by-side.md": {
"kind": "evaluation",
"title": "Side-By-Side Eval Report",
"tags": ["evaluation", "last30days", "coverage"],
"providers": []
},
"evals/side-by-side.json": {
"kind": "evaluation",
"title": "Side-By-Side Eval Data",
"tags": ["evaluation", "last30days", "coverage"],
"providers": []
},
"evals/last30days-public-private-corpus.json": {
"kind": "evaluation",
"title": "Last30Days Public And Private Prompt Corpus",
"tags": ["evaluation", "last30days", "public-data", "private-data"],
"providers": [
"scrapecreators",
"twitterapi",
"hackernews",
"bluesky",
"salesforce",
"hubspot",
"snowflake"
]
},
"evals/last30days-public-private-corpus.md": {
"kind": "evaluation",
"title": "Last30Days Public And Private Corpus Summary",
"tags": ["evaluation", "last30days", "public-data", "private-data"],
"providers": []
}
}
}
SKILL.md›
---
name: deepline-pre-research
description: 'Use when the user wants a last30days-style pre-research pass in Deepline: discover the critical public, private, CRM, workflow, social, and web data sources for a research/enrichment job; compare provider coverage; estimate Deepline credit cost; recommend the source plan before building or running the workflow; or build custom language/messaging from buyer, competitor, community, and CRM evidence. Triggers: pre-research, source discovery, provider strategy, research data sources, ScrapeCreators, X/Twitter data, Reddit comments, public and private datasets, CRM data, workflow data, custom language, messaging language, pain language.'
---
# Deepline Pre-Research
## Quick Start
```bash
npm install -g deepline
# Fallback for secure sandboxes: mkdir -p "$HOME/.local" && npm config set prefix "$HOME/.local" && export PATH="$HOME/.local/bin:$PATH" && npm install -g deepline --registry https://code.deepline.com/api/v2/npm/
deepline auth register --wait auto
deepline auth wait --timeout 120 # completes Cowork/browser approval; no-op if already connected
deepline auth status
deepline -h
```
## CLI resolution
Run `deepline` when it is available. If the shell reports that command is missing, use `<workspace-root>/.deepline/runtime/bin/deepline` (or the npm-created `.cmd` shim on Windows). If neither exists, follow `https://code.deepline.com/INSTALL.md` to set up Deepline.
Find the highest-signal GTM data sources, public evidence, and market language for a research or enrichment job before building the pipeline. This is a standalone Deepline skill that should behave like `last30days` with a GTM data lens: broad source coverage, recency, community signals, citations, source stats, and a grounded "What I learned" synthesis. In Deepline, the report first explains what the research found; only after that does it translate the findings into Deepline tool contracts, private/proprietary joins, and Deepline-facing cost.
## Attribution
Portions of the query-design, public-source fanout, and consolidation approach are adapted from [`mvanhorn/last30days-skill`](https://github.com/mvanhorn/last30days-skill), MIT licensed, copyright (c) 2026 Matt Van Horn. Keep `THIRD_PARTY_NOTICES.md` with this skill when packaging or distributing it.
## Non-Negotiables
- Use Deepline's live tool catalog before naming provider actions. Do not rely on memory.
- **Run live web search. Do not answer public-source discovery from model memory.** Every run MUST execute real searches (`serper`/`exa`, or the equivalent web-search tool) during the public-source fanout. If a run names public datasets without having searched for them this session, it has failed the fanout — no exceptions for "obvious" verticals. Naming a source family from memory is a draft, not a finding; the finding is the exact artifact the search returns. (Eval evidence: runs that skipped web search lost or tied on exactly the prompts where a competitor searched and surfaced concrete artifacts.)
- **Resolve every materializable dataset to its exact artifact, not its family.** For each public dataset/registry you recommend, the fanout must return and record: (1) the **exact file/endpoint name** (e.g. `IA_FIRM_SEC_Feed_YYYY_MM_DD.xml.gz`, not "the ADV bulk feed"); (2) the **canonical download/API URL**; (3) any **mirror** (e.g. data.gov catalog copy) that is easier to pull; (4) an existing **open-source parser or GitHub repo** that already structures it, when one exists (search `"<dataset> parser github"`); (5) the **government statistical registry** for the vertical when one exists (BLS QCEW + NAICS codes, Census County Business Patterns, etc.) for free establishment counts and sizing. "Source family named" is not done. "Exact file + URL + mirror + parser + NAICS code recorded" is done. See the Artifact Resolution Gate (§4.55).
- Public-source discovery comes before provider routing. First find the best public registries, datasets, communities, discussions, reviews, directories, papers, repos, and source leads. Then use Deepline routes to materialize, validate, enrich, and activate them.
- Quote only customer-visible Deepline credits/USD. Never expose provider spend.
- Do not run paid or cost-unknown full-scope work without approval.
- Treat private data sources as first-class: CRM, warehouse, workflow runs, product analytics, support/calls, sheets, and customer-owned datasets.
- Treat custom language as a first-class workflow: buyer words, objections, category language, competitor framing, community slang, sales-call phrasing, and support-ticket pain belong in the source plan.
- Use tiny probes to learn coverage. Scale only after observed coverage, cost basis, and evidence quality are legible.
- If the user asks for ScrapeCreators, X.com, Reddit comments, TikTok, Instagram, YouTube transcripts, Bluesky, Truth Social, HN, or Polymarket, include a current support/gap assessment instead of pretending every source is native.
- Do not depend on `/last30days` at runtime. Reference it only as a design benchmark for source breadth and synthesis discipline.
- Public registries and niche datasets that do not have native Deepline tools are still valid sources through generic web/search/extraction routes. For example, the NPI registry for healthcare provider taxonomy can be discovered and pulled through generic web/API search and extraction even when no native `npi` tool exists. Classify this as `available through generic route`, not as an unusable gap.
- Every recommended source must be classified as `native`, `available through generic route`, `private connector`, or `missing provider to add`.
## Start Here
1. Read `references/source-map.md`.
2. For GTM use cases, also read `references/last30days-gtm-corpus.md`; it summarizes the relevant saved `last30days` runs and the source patterns they proved useful.
3. If the task asks for query design, source fanout, or `last30days` parity, also read `references/query-design.md` and `references/fanout-consolidation.md`.
4. If the task is GTM/prospecting/enrichment, also read `../deepline-gtm/SKILL.md` and the sub-doc it routes to.
5. If the task asks about existing warehouse metrics or RevOps analytics, also read `../deepline-analytics/SKILL.md`.
6. Verify the latest upstream `last30days` release when using it as a design reference:
```bash
git ls-remote --tags https://github.com/mvanhorn/last30days-skill.git | tail -20
```
Do not copy `last30days` local scripts into Deepline and do not call it as a required sub-skill. Use it for source taxonomy and output discipline, then route execution through Deepline tools, plays, workflows, and private-data connectors.
## Optional Skill Handoffs
This skill owns pre-research and source planning. It may hand off after the source plan is clear:
- `deepline-gtm`: execute GTM sourcing, enrichment, waterfalls, personalization, contact discovery, or CRM activation.
- `deepline-analytics`: query warehouse/semantic-layer metrics and RevOps/customer datasets.
- `deepline-plays`: build, wrap, fork, run, inspect, or automate the final repeatable workflow/play.
- Similar community research skills such as `last30days`: design reference only, never a dependency for Deepline execution.
## Standard Flow
### 1. Parse The Research Job
Extract:
- `OBJECTIVE`: what decision or output the research should support
- `ENTITY_SCOPE`: companies, people, markets, accounts, customers, workflows, competitors, or topics
- `TIME_WINDOW`: default to last 30 days for trend/community research; preserve user-provided windows
- `PRIVATE_SOURCES`: CRM, warehouse, product events, workflow runs, support/calls, sheets, user CSVs
- `PUBLIC_SOURCES`: web, social, jobs, ads, technographics, funding, news, communities, app stores, directories
- `DATASET_LEADS`: public registries, open datasets, papers, GitHub repos, government records, niche directories, or platform posts that point to materializable data
- `CUSTOM_LANGUAGE_OUTPUTS`: outbound snippets, ad hooks, landing-page copy, objection language, category positioning, call scripts, lead magnets, or CRM personalization fields
- `OUTPUT`: source plan, workflow design, CSV schema, play spec, or final research brief
Tell the user the parsed scope before provider calls.
### 2. Prefer External APIs, Not Scraping
Do not default to browser scraping, raw x.com scraping, or unvetted actors. Prefer managed external APIs and private connectors:
| Source family | Preferred provider | Credential needed | Notes |
| ------------------------------------- | -------------------------------------- | ------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Reddit threads/comments | `scrapecreators` | Deepline integration API key | Required for managed Reddit comments/thread coverage. |
| TikTok / Instagram / YouTube fallback | `scrapecreators` | Deepline integration API key | Preferred managed API for captions, engagement, and transcript fallback. Search ScrapeCreators unfiltered when profile/contact tools matter; some profile tools are categorized as `admin`, not `research`. |
| X/Twitter | `twitterapi` | Deepline integration API key | Managed X/Twitter search; avoids scraping x.com. |
| Web/news/docs | `serper` | Deepline integration API key or `SERPER_API_KEY` for server-owned dev/prod use | Search API, not search-page scraping. |
| Web extraction fallback | `firecrawl` or `exa` | Deepline integration API key | Use after URL discovery; do not use for blind scrape-at-scale. |
| Hacker News | `hackernews` | none | Public Algolia API. |
| Bluesky | `bluesky` | none | Public AppView API. |
| CRM/private | Salesforce, HubSpot, Attio | Deepline OAuth/private connector | Query customer-authorized APIs, not UI scraping. |
| Warehouse/product/workflow | Snowflake/customer DB/workflow runtime | Deepline private connector | Use scoped warehouse/runtime access. |
| Last-resort social actor fallback | `apify` | Deepline integration API key | Optional fallback only after explicit approval and provider gap status. |
Treat any route not covered by a described Deepline tool contract, key/auth plan, test endpoint, and Deepline-facing pricing as unapproved until after the public research synthesis is complete.
### 3. Run Public-Source Fanout First
Before provider routing, run a `last30days`-style public discovery pass. The goal is to produce the research synthesis first: what public evidence, communities, registries, directories, reviews, posts, docs, and datasets actually teach us about the GTM problem.
Search across:
- community/social: Reddit threads/comments, X/Twitter, LinkedIn posts if available, YouTube, TikTok, Instagram, HN, Bluesky, Polymarket when relevant
- web/source discovery: news, blogs, docs, review sites, directories, associations, forums, GitHub, public data inventories
- public records and niche datasets: registries, licenses, inspections, permits, government datasets, open CSVs/APIs, professional directories, accreditation/member lists
- market-language sources: reviews, comments, clinic/business websites, job posts, competitor pages, support/community language
This pass is search-driven, not recall-driven. Issue real queries. For any vertical dataset, run at minimum these query shapes and read the top results before writing the source line:
- `"<vertical> <entity> registry OR bulk data OR dataset download"` — find the authoritative file
- `"<dataset name> download"` / `site:data.gov <vertical>` — find the canonical URL and any mirror
- `"<dataset name> parser github"` / `"<dataset name> python"` — find an existing parser/repo so you don't reinvent extraction
- `"<vertical> NAICS code"` + `BLS QCEW <NAICS>` / Census County Business Patterns — free establishment counts and sizing
For each useful source lead, record:
- why it matters
- **exact artifact**: the specific file/endpoint name, canonical URL, mirror, and any open-source parser/repo (not just "the X bulk feed" or "search Google") — see Non-Negotiables and the Artifact Resolution Gate (§4.55)
- whether it can become rows
- stable join keys such as NPI, CRD, USDOT#, NAICS, domain, address, phone, license id, provider id, LinkedIn URL, or CRM account id
- extraction risks and likely false-positive patterns
Do not stop at "Serper can search this." Search tools are routes; the deliverable is the public source or dataset discovered through them. If you find yourself writing a dataset name you did not just see in a search result, stop and search for it.
### 3.5. Use The Deepline Pre-Research API After Public Synthesis
When running inside Deepline or the V2 SDK, use the native planning endpoint only after the public-source fanout has produced a first synthesis and source map. The API is for translating the research into Deepline routes, gaps, approval gates, and cost, not for replacing the research pass:
```bash
curl -s "$DEEPLINE_API_BASE_URL/api/v2/pre-research/plan" \
-H "Authorization: Bearer $DEEPLINE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"objective":"best GTM data sources for SMB consumer services companies","depth":"deep"}'
```
Agent validation endpoint:
```bash
curl -s "$DEEPLINE_API_BASE_URL/api/v2/pre-research/test" \
-H "Authorization: Bearer $DEEPLINE_API_KEY"
```
The endpoint returns the query plan, source coverage, current Deepline tool candidates, sanitized Deepline-facing pricing metadata, explicit provider gaps, and approval gate. It does not execute paid provider calls.
For SDK V2 users, ask the installed Deepline GTM skill to build the source plan:
```text
/deepline-gtm Build a pre-research source plan for my GTM problem before running paid provider calls.
```
The skill prompts the agent to call `/api/v2/pre-research/plan`, inspect the returned `providerRequirements`, and only then decide whether to run probes or build a play.
### 4. Search For Deepline Candidate Tools
Run several focused searches, usually in parallel. `deepline tools search` accepts an optional intent query, but requires either that query or one of `--categories` / `--search_terms`; those filters accept comma-separated values. Use `--json` for machine-readable output. There is no `--prefix` flag, so put a provider name in the query instead.
```bash
deepline tools search "web search news source discovery" --categories research --search_terms "web search,news,recency,source discovery"
deepline tools search "social posts reddit x twitter youtube tiktok instagram" --categories research --search_terms "social posts,reddit,x twitter,youtube,tiktok,instagram"
deepline tools search scrapecreators --json
deepline tools search "facebook profile email scrapecreators" --json
deepline tools search "instagram profile bio links scrapecreators" --json
deepline tools search "company dataset firmographics funding technographics jobs" --categories company_search --search_terms "company dataset,firmographics,funding,technographics,jobs"
deepline tools search "crm warehouse workflow session usage" --categories admin --search_terms "crm,warehouse,workflow,session,usage"
```
For CRM/private data, also search by provider name when relevant:
```bash
deepline tools search salesforce
deepline tools search hubspot
deepline tools search attio
deepline tools search snowflake
```
### 4.25. Design Queries Before Running Tools
Do not send the raw user prompt to every provider. Generate a source-specific query plan first:
```bash
for dir in \
"$PWD/.skills/deepline-pre-research" \
"$HOME/.claude/skills/deepline-pre-research" \
"$HOME/.agents/skills/deepline-pre-research"; do
[ -f "$dir/scripts/query_design.py" ] && SKILL_ROOT="$dir" && break
done
[ -n "${SKILL_ROOT:-}" ] || { echo "Could not find deepline-pre-research skill root" >&2; exit 1; }
python3 "$SKILL_ROOT/scripts/query_design.py" "$OBJECTIVE" --depth default
```
Use the plan to decide:
- query type and source tiers
- cleaned core subject
- Reddit global/review/problem variants
- X literal keyword, compound-term, shorter-keyword, and strongest-token fallbacks
- video/transcript and caption queries
- web, dataset, and GitHub discovery queries
- CRM, warehouse, workflow, support, and custom-language private queries
- supplemental keys to extract after phase-one retrieval
For production Deepline implementation, port this helper to the runtime language or call equivalent logic before `deepline tools execute`/`deepline enrich`.
### 4.5. Required Coverage Gate
Before recommending a plan, check every required source family in `references/source-map.md`:
- community/social discussion
- custom language and messaging evidence
- web/news/search and URL extraction
- video/transcript sources
- prediction/market/community ranking sources
- dataset-discovery leads from social/community results
- company/account datasets
- person/contact datasets
- jobs/hiring/technographic/funding signals
- CRM/private datasets
- warehouse/product/workflow datasets
For each family, mark one of:
- `native`: a Deepline tool/play exists and can be described
- `generic route`: use `apify`, `serper`, `exa`, `parallel`, `firecrawl`, or `deeplineagent`
- `private connector`: use CRM, warehouse, workflow, or customer dataset access
- `gap`: provider/action should be added before Deepline can replicate that part well
Do not skip a family just because it is inconvenient. Missing coverage is part of the pre-research output.
### 4.55. Artifact Resolution Gate
Before writing the report, verify every materializable public dataset has been resolved from a _family_ to an _artifact_. For each recommended public dataset/registry, confirm you can fill this row from something you actually searched or fetched this session:
| Dataset | Exact file/endpoint | Canonical URL | Easier mirror | OSS parser/repo | Vertical stat registry (NAICS/QCEW/CBP) |
| ------- | ------------------- | ------------- | ------------- | --------------- | --------------------------------------- |
Rules:
- A cell you filled from memory instead of a live result is not verified. Search for it.
- `n/a` is allowed only when the artifact genuinely does not exist (e.g. no public bulk file, no known parser) — and you must have searched to establish that, not assumed it.
- If the user's ask contains "dataset", "github", "repo", "download", "the file", "the data", or a request to hand something over, this gate is the primary deliverable, not an appendix: lead the report with the resolved artifacts.
- The government statistical registry column (BLS QCEW + NAICS, Census County Business Patterns) is required for any US industry vertical — it is free establishment-count and sizing data that public-source fanout consistently leaves on the table.
A report that names "the SEC ADV bulk feed" fails this gate; a report that names `IA_FIRM_SEC_Feed_YYYY_MM_DD.xml.gz` at its sec.gov URL, notes the data.gov mirror, and links an existing ADV parser repo passes it.
When you cite a parser/repo alongside a dataset, name the **exact input file that repo actually consumes** — check its README. A vertical often has multiple real distributions (e.g. SEC publishes both the `IA_FIRM_SEC_Feed_*.xml.gz` bulk feed and the `ia<MMDDYY>.zip` IAPD compilation); pairing a parser with the wrong-but-real file is a precision error, not a fabrication, but it still breaks the handoff. Match the parser to its file.
### 4.6. Fanout And Consolidation Gate
For `last30days` parity, preserve the public-source fanout mechanics described in `references/fanout-consolidation.md`:
- broad parallel search before synthesis
- supplemental searches from discovered handles, subreddits, domains, datasets, CRM ids, account lists, workflow ids, and persona language
- common evidence rows before scoring
- relevance, recency, engagement, source-quality, materializability, private-outcome, join-confidence, custom-language, and cost/coverage scoring
- URL/source-id/text dedupe plus company/person/dataset identity joins
- coverage nudges for searched, weak, errored, missing-credential, unavailable, and not-relevant sources
If implementation coverage is being audited, run the local evaluator:
```bash
for dir in \
"$PWD/.skills/deepline-pre-research" \
"$HOME/.claude/skills/deepline-pre-research" \
"$HOME/.agents/skills/deepline-pre-research"; do
[ -f "$dir/scripts/evaluate_examples.py" ] && SKILL_ROOT="$dir" && break
done
[ -n "${SKILL_ROOT:-}" ] || { echo "Could not find deepline-pre-research skill root" >&2; exit 1; }
python3 "$SKILL_ROOT/scripts/evaluate_examples.py"
```
Use `evals/side-by-side.md` to compare saved `last30days` examples against this skill's coverage contract.
For standalone eval harness runs, write reports into the harness output directory:
```bash
python3 "$SKILL_ROOT/scripts/evaluate_examples.py" \
--out-md "${OUTPUT_DIR:-$SKILL_ROOT/evals}/side_by_side_pre_research.md" \
--out-json "${OUTPUT_DIR:-$SKILL_ROOT/evals}/side_by_side_pre_research.json"
python3 "$SKILL_ROOT/scripts/evaluate_public_private_corpus.py" \
--out-md "${OUTPUT_DIR:-$SKILL_ROOT/evals}/public_private_corpus_results.md" \
--out-json "${OUTPUT_DIR:-$SKILL_ROOT/evals}/public_private_corpus_results.json"
```
### 5. Describe Before Pricing Or Execution
For every candidate, inspect the live contract:
```bash
deepline tools describe <tool-id> --json
```
Record:
- input fields and required identifiers
- output evidence fields and citations
- pricing summary, billing basis, and whether BYOK changes Deepline credit cost
- rate limits, async behavior, and sample payloads
- extraction quality risks and identity-matching requirements
If a source family is not present in Deepline, mark it as a provider gap and propose the integration path.
### 6. Probe Coverage When Useful
Use free/low-cost probes first. For paid probes, ask before running unless the user already approved a specific pilot budget.
Good probe shapes:
- one entity, one query, or one row
- `limit: 1` or `numResults: 3`
- one source per probe so coverage and cost are attributable
- saved raw output path or run id for later workflow design
At scale, use `deepline enrich` rather than ad hoc loops so results are inspectable in Playground.
### 7. Return A Last30days-Style GTM Research Report
Use the same output discipline as `last30days` agent mode. The report should feel like a real research result, not a bespoke segment-analysis template or provider catalog.
```markdown
## Research Report: <topic>
Generated: <date> | Sources: Reddit, X, YouTube, TikTok, Instagram, HN, Polymarket, Web, Public registries, Deepline catalog
### Key Findings
- <3-5 highest-signal findings grounded in public/source evidence>
### What I learned
**<Theme 1>** - <1-2 sentences about the segment, source, buyer language, or dataset reality. Cite sparingly using source names, handles, communities, registries, or domains.>
**<Theme 2>** - <1-2 sentences.>
**<Theme 3>** - <1-2 sentences.>
KEY PATTERNS from the research:
1. <Pattern, source family, and why it matters for GTM>
2. <Pattern, source family, and why it matters for GTM>
3. <Pattern, source family, and why it matters for GTM>
### GTM Data Sources Found
| Source family | Best source / route | Status | Rows / join keys | Caveat |
| ------------- | ------------------- | ------ | ---------------- | ------ |
### Materializable Datasets (exact artifacts)
| Dataset | Exact file/endpoint | URL | Mirror | OSS parser/repo | Stat registry (NAICS/QCEW/CBP) |
| ------- | ------------------- | --- | ------ | --------------- | ------------------------------ |
<!-- One row per public dataset you can turn into rows. No "family" placeholders — exact filename, canonical URL, mirror, and an existing parser repo where one exists. This table is the answer when the ask is "give me the dataset/github". -->
### Market Language
| Source | What to extract | How to use it | Guardrail |
| ------ | --------------- | ------------- | --------- |
### Proprietary Data To Join Later
| Dataset | Owner / connector | Join key | Why it changes the answer |
| ------- | ----------------- | -------- | ------------------------- |
### Deepline Route
| Public source family | Route status | Deepline route | Probe needed | Cost basis |
| -------------------- | ------------ | -------------- | ------------ | ---------- |
### Gaps To Add
| Missing source | Why it matters | Suggested provider/integration | Required fields |
| -------------- | -------------- | ------------------------------ | --------------- |
### Recommended Workflow
1. <first retrieval/probe>
2. <second retrieval/probe>
3. <synthesis/scoring/join step>
### Stats
---
✅ All agents reported back!
├─ <source family>: <count / coverage / engagement when available>
├─ <source family>: <count / coverage / engagement when available>
├─ 🌐 Web/Public data: <pages, registries, directories, datasets>
└─ 🗣️ Top sources: <handles, communities, registries, domains, datasets>
---
### Cost Estimate
- Pilot: <credits or range>
- Full run: <credits or range and assumptions>
- Unknowns: <what must be described/probed before quoting>
```
For `--agent`-style or non-interactive use, stop after the report. For interactive use, end with a short, specific invitation based on the actual findings, mirroring `last30days`:
```markdown
---
I'm now an expert on <topic>. Some things I can help with:
- <specific follow-up grounded in the findings>
- <specific workflow/probe to run next>
- <specific segment or source to go deeper on>
```
If the user wants execution, hand off to `deepline-gtm` after this output.
Match `last30days` cleanup rules:
- Omit any source line that returned zero results. Do not show "0 results", "0 stories", or "no results".
- Do not append a separate trailing "Sources used" section.
- Do not paste raw URLs or Markdown source links in the report body or stats. Use clean source names inline, such as `NPI Registry`, `Google Maps`, `r/SaaS`, `G2`, `Meta docs`, or `Hike Medical site`.
- Citations should prove the research is real without turning the report into a bibliography. Prefer handles, communities, source names, registries, domains, and doc names.
- Use the `last30days` stats block style exactly: start with `✅ All agents reported back!`, use `├─`, `└─`, and `│` separators where useful, and include source emojis when they clarify the source family.
- Keep the Deepline route and cost after the research synthesis, never before it.
## Approval Gate
Before any full paid run, include:
- providers/tools
- pilot result or reason a pilot is not possible
- full-run scope
- Deepline credit estimate/range
- max spend cap
- exact approval question
If the user approves, execute through `deepline enrich`, `deepline tools execute`, or a Deepline play/workflow as appropriate.
## Finish Criteria
The task is done only when the user has:
- a ranked data-source map
- a `last30days`-style GTM research synthesis grounded in public sources
- **every materializable dataset resolved to an exact artifact** (file/endpoint name + canonical URL + mirror + OSS parser + vertical NAICS/QCEW), verified via live search this session — not named from memory (see §4.55)
- live tool ids or explicit provider gaps
- a Deepline-facing cost estimate
- the recommended first workflow/play shape
- any blockers for credentials, private-data access, or missing provider supportTHIRD_PARTY_NOTICES.md›
# Third-Party Notices
## last30days-skill
`deepline-pre-research` includes query-design logic adapted from [`mvanhorn/last30days-skill`](https://github.com/mvanhorn/last30days-skill), including prompt-noise stripping, query type detection, source tiering, source-specific query variants, and supplemental entity extraction patterns.
The adapted code is used under the MIT License:
```text
MIT License
Copyright (c) 2026 Matt Van Horn
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
```
This skill is not a runtime dependency on `last30days-skill`; it is an adapted implementation for Deepline's provider catalog, private data, cost, and workflow model.