SKILL DETAIL
lookalike-audience-seeds
mbfinotti/advertising-skills/lookalike-audience-seeds
Select and size the seed customer list behind a lookalike, similar, or value-based audience - which customers to upload, how many, and whether the matched count (not the row count) clears the platform floor - with RFM and value-based selection, a negative-selection pass, a privacy and consent gate before any customer list upload, and a fallback ladder when the floor cannot be cleared. Use whenever the user mentions a lookalike or similar audience, a seed list, a customer list upload, match rate, an audience that is too small, or says their lookalike isn't working - even if they never say 'seed'. Covers B2B and B2C. Do NOT use to design the whole targeting tier structure - use mbfinotti/advertising-skills@ad-audience-targeting instead.
Installation
npx skills add https://github.com/mbfinotti/advertising-skills --skill lookalike-audience-seeds
技能檔案
SKILL.md
最近同步 · 2026年9月24日
evals/evals.json›
{
"skill_name": "lookalike-audience-seeds",
"evals": [
{
"id": 1,
"prompt": "We're Corvid Analytics, a B2B data-quality platform with an ACV around $42K. I exported our closed-won contacts from the CRM - 700 people, work emails only - and I want to build a LinkedIn lookalike from them to find similar buyers. 700 is comfortably above LinkedIn's 300 minimum so we should be fine, right? If we need more volume I can always throw in our MQL list, there are about 6,000 of those. Walk me through how you'd set this up.",
"expected_output": "A corrected plan that computes the effective matched seed, shows the 700-contact list fails LinkedIn's 300 matched floor, refuses MQL padding, and routes through identity enrichment (fallback rung 1) with its compliance cost, targeting Predictive Audiences rather than the sunset Lookalike product.",
"files": [],
"expectations": [
"States that LinkedIn's 300 minimum applies to matched members, not uploaded rows, and rejects the user's row-count reading",
"Applies an expected match rate around 30-40% for a work-email contact list on LinkedIn",
"Computes the effective seed at roughly 210-280 matched (700 x 0.30-0.40) and concludes it fails the 300 matched floor",
"Corrects the 'LinkedIn lookalike' framing: LinkedIn Lookalike Audiences were sunset in February 2024, replaced by Predictive Audiences",
"Refuses to add the 6,000 MQLs to the seed, keeping closed-won as the selection and stating that selection quality is never the first thing traded (SQL-only beats all-MQL)",
"Recommends fallback ladder rung 1, identity enrichment appending personal identifiers, as the route to the floor",
"Explains why the cheap rungs are structurally empty here: an all-time closed-won list has no recency window left to widen and no adjacent segment that is not lost deals or MQLs",
"Mentions expanding won accounts into contacts via title, seniority, and function filters before upload",
"Notes that enrichment brings a vendor DPA, counsel or lawful-basis re-check, and roughly a week or more of procurement coordination",
"Cites LinkedIn sizing guidance of roughly 10,000+ uploaded emails for a contact list (or 1,000+ companies for a company list) to reliably clear 300 matched",
"Advises judging early results on match quality and engagement only, deferring the seed verdict to cohort maturity because the B2B sales cycle lags the conversion read",
"Recommends sizing the post-enrichment upload at 2-3x the target matched count",
"States that if no procurement route or vendor DPA exists, rung 1 is deleted from the ladder and the honest finding is that no ladder clears the floor - reported as the finding rather than filled with a looser seed"
]
},
{
"id": 2,
"prompt": "I run growth at Lumme Botanica, a DTC skincare brand with about 62,000 lifetime customers. We want a value-based lookalike on Meta. My plan: take the top 15% of customers by total lifetime spend and upload that with lifetime spend as the value column. Our orders are split between our US store (USD) and EU store (EUR), so I'd merge both into one file. Some rows are gift recipients with $0 spend, and one corporate account has spent about 200x our median customer. Sound good?",
"expected_output": "A rejection of cumulative lifetime spend as the value column with margin or predicted LTV substituted, fixes for the zero values, mixed currencies, and whale outlier, upload-mapping and terms-acceptance checks, and a proper RFM or value-slice selection with recency window and negative-selection pass.",
"files": [],
"expectations": [
"Rejects cumulative lifetime spend/revenue as the value column",
"Explains the reason: cumulative revenue tracks tenure, and high-revenue low-margin customers poison the model",
"Recommends margin or predicted LTV as the value column instead",
"Flags the $0 gift-recipient rows: the value column must contain positive values only",
"Flags the USD/EUR merge: one currency per list, since platforms do not normalize mixed currencies",
"Recommends capping or winsorizing the 200x corporate whale so a few large accounts do not skew the model",
"Requires refunded amounts to be excluded from the value computation",
"Requires confirming the value column is actually mapped at upload",
"Warns that an unmapped value column silently builds a standard (not value-based) lookalike, and that the value-based option being absent at audience-build time means the mapping failed",
"Notes that value-based audiences on Meta require a separate terms acceptance",
"Recommends selection via RFM Champions + Loyal cohorts, or the top 10-25% by predicted LTV or margin, rather than the top 15% by spend",
"Applies a recency window (30-90 days default, up to 180 for pixel-tracked buyers) instead of the full lifetime list",
"Runs the negative-selection pass: strip refunders, chargebacks, serial returners, discount-only buyers, employees, and wholesale accounts",
"Keeps existing customers in the seed while suppressing them from the acquisition campaign at delivery level"
]
},
{
"id": 3,
"prompt": "Help me figure out our Meta match rate at Fenwright Home. We uploaded a 22,000-row customer CSV through the UI and matched only about 9%. Our security team insisted on SHA-256 hashing every email with a per-customer salt before the file left our network - glad we did that at least. Phones are stored like '+44 07911 123456', emails are mixed case with some trailing spaces, and about a third of the rows date from 2021-2022. Should we just buy a data enrichment service to fix the match rate?",
"expected_output": "A diagnosis naming the salted pre-hash as the primary cause, a normalization-first fix plan (plaintext UI upload, E.164-style phones, lowercased emails, more identifiers, stale-row cut) ordered ahead of enrichment, with the effective seed recomputed and match rate kept as a hygiene indicator only.",
"files": [],
"expectations": [
"Identifies the salted pre-hash as the primary cause: the platform hashes on ingest, so a salted or double hash never matches",
"Recommends re-uploading normalized plaintext through the UI, with hashing reserved for API or warehouse syncs that require it",
"Corrects the phone format to digits with country code, no plus sign, no leading zero (E.164-style, e.g. 447911123456)",
"Corrects emails to lowercased and trimmed",
"Recommends adding more identifier types per row, since multiple identifiers per row beat email alone",
"Flags rows older than 12-18 months as past the decay cliff and recommends dropping or archiving them",
"Warns that dropping old rows lifts the rate partly by shrinking the denominator, so the effective seed must still clear the floor afterwards",
"Orders the fixes as: normalize formats first, then join identifiers already held, then drop stale rows, with enrichment last",
"Answers the enrichment question with no: it is the lowest-ratio lever here and not the first move, reserved for cases like B2B work-email lists on consumer platforms",
"References Meta's documented band (50-80% on good lists, below 40% signaling a poor or outdated list) as the calibration reference",
"Mentions Meta's Event Match Quality score with the 6.0+ practitioner target",
"Recomputes or instructs recomputing the effective seed (rows x match rate) against the platform floor after the fixes",
"Does not treat the improved match rate as evidence of audience quality - match rate is a data-hygiene leading indicator only"
]
},
{
"id": 4,
"prompt": "We're Hafenlicht, a Berlin-based supplement brand. I want a Customer Match list in Google and a Meta custom audience from our 30,000 EU customers this week, then lookalikes off both. One segment idea: buyers of our sleep-and-anxiety product line converted best, so we'd seed from them. Our agency says consent isn't an issue because everything gets SHA-256 hashed before upload, so it's anonymous data and GDPR doesn't really apply. Legal review is backed up for a month - can we define and upload the seed now and let them review in parallel?",
"expected_output": "A hard stop at the privacy and consent gate: the hashing-equals-anonymous claim rejected with the German precedent, Consent Mode v2 required for EEA Customer Match, the health-category segment refused, upload-in-parallel refused, and platform-native engagement audiences named as the only currently available seed class.",
"files": [],
"expectations": [
"Stops at the privacy and consent gate before doing any selection work, rather than proceeding with the seed definition",
"Rejects the agency's claim: SHA-256 output is pseudonymous, not anonymous, so consent obligations survive hashing",
"References the German precedent - the Bavarian DPA position upheld by the Higher Administrative Court Munich (2018) - that customer-list custom audiences require prior consent",
"Requires Consent Mode v2 signals with both consent fields GRANTED for EEA use of Google Customer Match, mandatory since March 2024",
"Rejects the sleep-and-anxiety product segment as a seed built on a sensitive health category",
"Refuses to upload in parallel with legal review: no upload before the consent basis is confirmed",
"Names the specific missing gate items (documented lawful basis or consent, Consent Mode v2 status, platform terms acceptance) instead of a generic 'check with legal'",
"States that suppression, opt-outs, and deletions must be enforced once in the warehouse before any sync, not per-platform afterwards",
"Disclaims legal advice and routes the lawful-basis question to the user's counsel",
"Does not emit a Seed Specification, because an unconfirmed consent basis fails the evidence gate",
"Treats list-based seed sources as unavailable (deleted from the menu), not merely delayed",
"Offers platform-native engagement audiences as the seed class still available now, since they never leave the ad account and trigger no consent gate",
"Mentions platform customer-list terms acceptance for the specific ad account as a gate item"
]
},
{
"id": 5,
"prompt": "Bramblewick makes premium dog gear. Meta's docs say a good seed is 1,000-5,000 people, but we only have 850 buyers from the last 90 days. My plan: top up with our 24,000-person newsletter list plus everyone who opened a marketing email in the past 12 months - that gets us way over 5,000. We haven't signed any data vendors and don't plan to. Good approach?",
"expected_output": "A refusal to pad the seed, showing the 850 buyers already clear Meta's matched floor, with openers rejected in favor of clickers, and any extra volume sought through the fallback ladder - recency widening first given no enrichment vendor - never through undifferentiated subscribers.",
"files": [],
"expectations": [
"Refuses to pad the seed with newsletter subscribers to hit the recommended size",
"Cites the homogeneity principle: seed homogeneity affects audience effectiveness more than its size, per Meta's own guidance",
"Notes the practitioner consensus that a few hundred high-value customers outperform thousands of undifferentiated subscribers, labeled directional rather than audited",
"Distinguishes Meta's 100 matched floor from the 1,000-5,000 recommended band",
"Computes the effective seed: 850 rows at Meta's 50-80% band is roughly 425-680 matched, which clears the 100 floor",
"Rejects email openers as a source because Apple Mail Privacy Protection (2021) inflated opens into noise, allowing clickers only",
"If more volume is wanted, walks the fallback ladder instead of loosening selection quality",
"Applies the documented ladder flip (rung 2 before rung 1) because no enrichment vendor is contracted and there is recency left to widen: 90 to 180 days first",
"Mentions stacking an adjacent segment of the same homogeneity (rung 3) as a later step, not the newsletter list",
"If email clickers are ever exported, requires checking that ESP consent scope covers ad-platform upload rather than assuming it",
"Runs the negative-selection pass on the 850 buyers before export",
"Keeps existing customers in the seed and suppresses them from the acquisition campaign at delivery level",
"Recommends sizing the upload at 2-3x the matched target"
]
},
{
"id": 6,
"prompt": "Nordvik Living sells furniture and lighting online in Germany, France, the Netherlands, the UK, and Sweden - about 18,000 customers total, though Sweden is only around 120 of them. I was going to upload one combined customer list to Meta and build one lookalike targeting all five countries. I'd also love splits by product line (furniture vs lighting), a high-AOV split, and maybe language versions. It's just me managing this, updating lists by hand whenever I remember, and opt-outs get processed inside each ad platform separately when someone complains.",
"expected_output": "A per-country seed-and-lookalike plan replacing the single global audience, the split order applied with each split floor-checked (Sweden fails), split count capped to the single owner's capacity, a real refresh cadence with warehouse-first opt-out suppression, and an overlap check.",
"files": [],
"expectations": [
"States that Meta lookalikes are country-scoped by construction, so one global seed produces an averaged audience serving no market well",
"Requires one seed-and-lookalike pair per country",
"Applies the split order country > product line > high-AOV vs all, with language/region last because it usually correlates with country and adds the least new information",
"Checks each split's own effective seed against the floor before committing to it",
"Identifies the Sweden split as failing: roughly 120 customers at 50-80% match is about 60-96 matched, under Meta's 100 floor, so Sweden needs the fallback ladder or gets deferred",
"Caps the number of splits at what the single named refresh owner can maintain, advising against taking every proposed split at once",
"Notes each split carries a permanent standing cost: an audience to refresh, monitor for overlap, and keep above its floor forever",
"Notes that setup effort per split is identical (one WHERE clause) - the standing cost is what limits the count, not the query work",
"Rejects the 'whenever I remember' refresh: static CSV uploads do not auto-refresh, so a weekly-to-monthly cadence with a named owner is required",
"Recommends an automated warehouse sync over manual CSV on efficiency and compliance-cost grounds",
"Flags the per-platform opt-out handling: suppression must be enforced once in the warehouse before any sync, and a manual CSV re-uploads yesterday's opt-outs every cycle",
"Positions a CSV only as a bridge, with a named owner and a date the sync lands",
"Recommends an overlap check between the resulting audiences with the roughly 20-30% concern threshold",
"Frames the high-AOV split as a test (one comparison, keep the winner) run only after the country and product-line splits are settled"
]
},
{
"id": 7,
"prompt": "Perch & Pine, a home-decor DTC brand, launched a Meta lookalike from our customer list 5 days ago. Match rate came back 72%, so the seed is clearly solid, but CPA is running 40% over target already - I think the seed is bad and want to swap it tomorrow. Context: our previous seed version had amazing CPA, but finance later found those customers barely reordered and mostly bought with our 40%-off welcome code. We've never actually measured cohort LTV for acquired customers. Which seed do we keep? I was going to compare the two audiences head-to-head as two ad sets next week.",
"expected_output": "A measurement correction: match rate rejected as a seed grade, the 5-day read refused pending learning-phase exit, the prior version diagnosed as the low-CPA/weak-LTV pitfall with a re-seed on margin, cohort LTV by seed version installed as the KPI with a baseline-building first checkpoint, and the ad-set head-to-head replaced by a split test or geo holdout.",
"files": [],
"expectations": [
"Rejects '72% match so the seed is solid': match rate is a data-quality leading indicator only, never a grade of seed quality",
"Refuses to judge the new seed at 5 days because the learning phase has not completed",
"Cites roughly 50 optimization events per ad set per week, or 2-4 weeks at low volume, as the reading threshold",
"Diagnoses the previous seed version as the documented low-CPA/weak-LTV pitfall: the seed scaled low-margin converters",
"Prescribes re-seeding on margin or predicted LTV and retiring that seed version",
"Flags discount-only buyers (the 40%-off welcome-code cohort) as a negative-selection exclusion for the next seed",
"Names cohort LTV at 90/180/365 days by seed version as the KPI to grade seeds on",
"States the pass threshold: the seed version's 90-day acquired-customer cohort LTV meets or beats the account average at equal-or-better CPA",
"Given no measured baseline, states the seed cannot be graded yet and sets the first cohort checkpoint as the baseline-building run",
"Rejects the naive two-ad-set head-to-head comparison because audience overlap contaminates it",
"Recommends a user-level split test or a geo holdout as the valid comparison methods",
"Defaults to the split test, escalating to the geo holdout when the decision is expensive to reverse or the split test's own read is disputed",
"Notes the geo holdout's evidence advantage: it survives cross-device identity loss and measures incremental sales rather than platform-attributed ones",
"Treats the seed as a versioned artifact with a re-check scheduled one cohort window out"
]
},
{
"id": 8,
"prompt": "Our new agency sent Calloway Cycles (bike accessories, B2C plus a wholesale line) a targeting doc that looks recycled: (1) build Google Similar Audiences from our converter list; (2) LinkedIn Lookalike Audiences from our 2,400 wholesale-buyer contacts; (3) a TikTok lookalike from our 400 best customers - the doc says TikTok's minimum is 100, so that's fine; (4) a Google Demand Gen lookalike, the doc says 100 minimum there too; (5) on Meta, hard-restrict delivery to the lookalike so no budget leaks to anyone outside it. Anything we should push back on before signing off?",
"expected_output": "A point-by-point rebuttal: Similar Audiences and LinkedIn Lookalikes are dead products with named replacements, TikTok and Demand Gen floors corrected to the stricter 1,000 figures (400 rows cannot reach TikTok's floor), the Meta inclusion-lock corrected to suggestion-not-rule, and floors flagged as moving targets to verify live.",
"files": [],
"expectations": [
"Flags Google Similar Audiences as removed in August 2023, so item 1 is a dead product",
"Notes the replacement mechanism: first-party lists now act as signals (Customer Match) feeding optimized targeting and Demand Gen lookalikes",
"Flags LinkedIn Lookalike Audiences as sunset in February 2024",
"Names LinkedIn Predictive Audiences with a 300 matched minimum as the replacement",
"Mentions the cap of 30 Predictive Audiences per ad account",
"Takes TikTok's stricter 1,000 floor over the 100 figure, applying the rule that when a platform's own pages disagree you plan against the stricter number",
"Concludes that 400 customers cannot reach TikTok's 1,000 matched floor at any match rate, since uploaded rows are below the floor itself",
"Mentions the 10,000 practitioner-recommended TikTok seed as the realistic bar",
"Takes 1,000 active matched (API docs) as the Google Demand Gen lookalike floor over the Help Center's 100",
"Corrects the Meta hard-restrict belief: inclusion audiences behave as suggestions under Advantage+, and the delivery algorithm may serve outside them",
"States that only exclusion/suppression audiences remain hard rules",
"Notes Meta's detailed-targeting exclusions were removed in March 2025, leaving custom-audience exclusions under Audience Controls as the remaining hard suppression",
"Computes the LinkedIn wholesale list's effective seed: 2,400 contacts at roughly 30-40% work-email match is about 720-960 matched, clearing the 300 floor",
"Recommends verifying floors against the platforms' live documentation because floors move, citing Google Customer Match's cut from 1,000 to 100 in 2025 as the example"
]
},
{
"id": 9,
"prompt": "Marlowe Threads is a DTC apparel brand. The founder wants a prospecting lookalike live on Meta by Friday - hard date, board meeting. Problem: our customer data is scattered across three systems, no unified export exists, and our only data engineer is booked out for five weeks. A consultant told us 'purchasers are always the strongest seed, everything else is a downgrade - build the export first, launch later.' The Meta pixel has been on the site for a year with add-to-cart and checkout events firing, and our product videos get decent view counts. What would you actually ship by Friday?",
"expected_output": "A Friday-shippable plan built on platform-native engagement seeds (high-intent viewers or checkout initiators), the purchaser ranking's export-exists condition stated, the deadline named as what re-ranked the menu, quality floor held above page-engagers/all-visitors, and the starved purchaser-export asset scheduled as the compounding follow-up.",
"files": [],
"expectations": [
"Does not delay the launch to build the purchaser export first",
"States the condition behind the consultant's rule: purchasers lead the ranking when the customer export already exists, and here it does not",
"Recommends platform-native engagement seeds for Friday: high-intent page viewers and/or checkout/add-to-cart initiators",
"Names the hard deadline and data-team unavailability as the interview answers that re-ranked the source menu",
"Notes platform-native audiences build inside the ad account in minutes and skip the customer-list upload and its consent gate entirely",
"Refuses to drop below high-intent viewers/lead-form submitters on the ranking - no page-engagers or all-visitors seed, because past that point per-row quality falls faster than effort",
"If video viewers are used, requires the 75%+ completion tier rather than 3-second views",
"Flags the one judgment call for high-intent viewers: agreeing which pages count as intent",
"Notes engagement-sourced audiences auto-refresh (Meta roughly every 3-7 days), removing the manual refresh burden",
"Names what the efficiency order starves - the customer-identity asset (purchaser export / warehouse sync) - and schedules it as the compounding follow-up rather than dropping it",
"States a promote-it trigger for the export: a compounding-asset mandate, or the same account losing this round again",
"Notes platform-native audiences are durable against mobile-OS tracking loss, with retention windows up to 365 days",
"Respects the learning phase after the Friday launch: no performance verdict before roughly 50 optimization events per ad set per week or 2-4 weeks"
]
},
{
"id": 10,
"prompt": "Vellum Coffee, US-only DTC. We're ready for a Meta lookalike seed and I want the final spec today. Data: 21,000 customers; email plus phone on about 80% of rows; consent for ad matching captured at checkout; opt-outs sync to our warehouse nightly. Our BI tool has a column called 'ltv_score' someone built years ago - nobody remembers whether it's revenue, margin, or a model output, but let's use it as the value column since it's there. After filtering to purchasers from the last 120 days minus the usual junk (refunders, staff, two wholesale accounts), we land at 6,200 rows. For the match rate just use 95% so the numbers look safe. Target: new-customer CPA under $28; our account's 90-day acquired-customer LTV averages $54. Write the spec.",
"expected_output": "A Seed Specification built on the documented Meta match band instead of the invented 95%, with the unknown-provenance ltv_score refused as a value column (flat seed or provenance requested), effective seed computed against the floor, and consent, refresh, measurement, and re-check lines filled from the given data.",
"files": [],
"expectations": [
"Refuses the 95% match assumption and estimates from Meta's documented 50-80% good-list band, explicitly labeling it an estimate",
"Applies the evidence gate to 'ltv_score': a value column of unknown provenance means the seed is not value-based until someone can say what the column measures",
"Does not place 'ltv_score' in the spec as the value column - either specifies a flat seed or requests provenance first",
"Emits a Seed Specification block carrying the named fields: platform/audience type, definition, value column, row count, expected match, effective seed, fallback used, exclusions, consent basis, refresh, measurement, re-check",
"Computes the effective-seed line as rows x match vs floor: 6,200 x 0.50-0.80 is roughly 3,100-4,960 matched against Meta's 100 floor, so it clears (within the 1,000-5,000 recommended band)",
"States the expected-match basis on its line: platform documented range plus a list-quality adjustment",
"Fills the consent-basis line with the checkout consent and the nightly warehouse-enforced opt-out suppression",
"Fills the refresh line with a cadence, a named owner, and the mechanism (manual CSV vs automated sync)",
"Fills the measurement line with the $28 new-customer CPA target and cohort-LTV checkpoints by seed version against the $54 account average",
"Sets the re-check date one cohort window out (roughly 90 days)",
"Records the negative-selection pass in the exclusions line and delivers existing-customer suppression as a separate delivery-level audience, keeping purchasers in the seed",
"Addresses the 120-day recency window against the guidance: acceptable only for pixel/CAPI-tracked buyers (default 30-90 days), otherwise tightened toward 90",
"States the pass threshold: the seed version's 90-day acquired-customer cohort LTV must meet or beat the $54 account average at equal-or-better CPA"
]
}
],
"trigger_queries": [
{ "query": "Which customers should I upload to build a Meta lookalike?", "should_trigger": true },
{ "query": "My lookalike audience isn't performing - could the source list be the problem?", "should_trigger": true },
{ "query": "How many people do I need in a customer list for a LinkedIn predictive audience?", "should_trigger": true },
{ "query": "Our custom audience match rate is only 12%, what's wrong with our file?", "should_trigger": true },
{ "query": "TikTok says my audience is too small to serve, what now?", "should_trigger": true },
{ "query": "Should I use my newsletter list as the source for a similar audience?", "should_trigger": true },
{ "query": "Help me pick the seed for a value-based lookalike", "should_trigger": true },
{ "query": "What's the minimum audience size for Customer Match lookalikes these days?", "should_trigger": true },
{ "query": "Can I build a lookalike from only 400 customers?", "should_trigger": true },
{ "query": "Best customers to seed a lookalike from - top spenders or most recent buyers?", "should_trigger": true },
{ "query": "I uploaded 8,000 emails to Meta but the audience won't build", "should_trigger": true },
{ "query": "Do I need to hash my customer CSV before uploading it to an ad platform?", "should_trigger": true },
{ "query": "which crm segment should feed our similar audience on google", "should_trigger": true },
{ "query": "We're a B2B SaaS with 350 closed-won customers - can we do lookalikes at all?", "should_trigger": true },
{ "query": "How recent should the customers in my seed list be?", "should_trigger": true },
{ "query": "Should refunders be removed from our customer list before we upload it for ads?", "should_trigger": true },
{ "query": "Is a 30% match rate on our uploaded list normal?", "should_trigger": true },
{ "query": "How do I make a value-based audience actually use our LTV data?", "should_trigger": true },
{ "query": "Our lookalike went stale - do customer lists need refreshing?", "should_trigger": true },
{ "query": "One lookalike for all our European markets, or one per country?", "should_trigger": true },
{ "query": "What list size do I need before Meta's lookalike gets good?", "should_trigger": true },
{ "query": "The audience we uploaded shows fewer people than the rows in the file - why?", "should_trigger": true },
{ "query": "Seed list strategy for a new prospecting audience", "should_trigger": true },
{ "query": "Can I use email subscribers who opened our campaigns as an ad audience source?", "should_trigger": true },
{ "query": "Are work emails okay for a TikTok custom audience upload?", "should_trigger": true },
{ "query": "GDPR and uploading customer lists to Facebook - what do we need before we do it?", "should_trigger": true },
{ "query": "My similar audience on Google disappeared, what replaced it?", "should_trigger": true },
{ "query": "What percentage of uploaded contacts typically match on LinkedIn?", "should_trigger": true },
{ "query": "Should our seed be our best 500 customers or all 20,000?", "should_trigger": true },
{ "query": "how big should a lookalike source audience be", "should_trigger": true },
{ "query": "We only have 900 buyers - pad the audience with site visitors?", "should_trigger": true },
{ "query": "Setting up a predictive audience from our contact list, walk me through sizing", "should_trigger": true },
{ "query": "Is it bad to seed a lookalike from everyone who ever bought, even 5 years ago?", "should_trigger": true },
{ "query": "Customer list upload keeps getting rejected as too small", "should_trigger": true },
{ "query": "What identifiers should each row of an ad-platform customer upload have?", "should_trigger": true },
{ "query": "Do lookalikes work with a phone-number-only list?", "should_trigger": true },
{ "query": "Value column for a Meta value-based audience - revenue or margin?", "should_trigger": true },
{ "query": "How often should we re-upload our customer list to ad platforms?", "should_trigger": true },
{ "query": "Building audiences from our warehouse - what should the export contain for a lookalike?", "should_trigger": true },
{ "query": "Why does our lookalike convert cheap but the customers never come back?", "should_trigger": true },
{ "query": "Our agency wants to build ad audiences from our anxiety-supplement buyers - any issue?", "should_trigger": true },
{ "query": "similar audience seed size for snapchat?", "should_trigger": true },
{ "query": "What counts as a good source audience for Advantage+ lookalikes?", "should_trigger": true },
{ "query": "The ads platform says minimum 1,000 matched - our list is 3,000 rows, are we fine?", "should_trigger": true },
{ "query": "Which is better as an audience source: checkout abandoners or video viewers?", "should_trigger": true },
{ "query": "Should I exclude current customers from the list I upload for prospecting lookalikes?", "should_trigger": true },
{ "query": "First lookalike for our Shopify store - where do I start with the customer list?", "should_trigger": true },
{ "query": "Does list quality really matter more than list size for lookalikes?", "should_trigger": true },
{ "query": "Audience match rate dropped after we salted the hashes - related?", "should_trigger": true },
{ "query": "Pinterest actalike source list requirements?", "should_trigger": true },
{ "query": "We want to find people similar to our top 10% of customers - how do we define that 10%?", "should_trigger": true },
{ "query": "Can a 600-contact ABM list power a LinkedIn audience?", "should_trigger": true },
{ "query": "whats the deal with uploading EU customer emails to google ads now", "should_trigger": true },
{ "query": "How do I know if my uploaded audience is big enough before spending?", "should_trigger": true },
{ "query": "Best practices for the customer file behind a Customer Match campaign", "should_trigger": true },
{ "query": "Our lookalike audience never left 'populating' - is my list the issue?", "should_trigger": true },
{ "query": "Do I split my seed list by product line or keep one big list?", "should_trigger": true },
{ "query": "Fixing a low match rate: enrichment vendor or clean the data first?", "should_trigger": true },
{ "query": "What's a realistic match rate for a 2-year-old email list?", "should_trigger": true },
{ "query": "Reddit ads customer list minimum - is our 5,000 list enough?", "should_trigger": true },
{ "query": "Grade my plan: seed the lookalike with everyone who opened an email in 12 months", "should_trigger": true },
{ "query": "I need an audience of people like our subscribers - what data do I feed the platform?", "should_trigger": true },
{ "query": "The board wants lookalikes live Friday but our customer export doesn't exist yet", "should_trigger": true },
{ "query": "Amazon DSP lookalike from our buyer list - how many user IDs do we need?", "should_trigger": true },
{ "query": "Are lookalike seeds supposed to be refreshed, or is upload once enough?", "should_trigger": true },
{ "query": "help, uploaded customer audience too small to use for a lookalike", "should_trigger": true },
{ "query": "Design our full paid targeting plan - cold, interest, lookalike, and retargeting tiers", "should_trigger": false },
{ "query": "How should we layer prospecting vs retargeting audiences across the funnel?", "should_trigger": false },
{ "query": "Set recency windows and frequency caps for our cart-abandoner retargeting sequence", "should_trigger": false },
{ "query": "Build a message ladder for each stage of our retargeting funnel", "should_trigger": false },
{ "query": "Map the buying committee for our $80K ACV deal and how to reach each role", "should_trigger": false },
{ "query": "Which job functions should we target for enterprise IT purchases?", "should_trigger": false },
{ "query": "Our Meta pixel is double-counting purchases - help me debug deduplication", "should_trigger": false },
{ "query": "Verify our conversion events fire once before the campaign launches", "should_trigger": false },
{ "query": "Meta reports 210 conversions, GA4 shows 92 - reconcile the gap", "should_trigger": false },
{ "query": "Why does our CRM show half the revenue the ad platform claims?", "should_trigger": false },
{ "query": "Our whole account's ROAS dropped 35% this month - diagnose it", "should_trigger": false },
{ "query": "CPA doubled across every campaign, where do I start?", "should_trigger": false },
{ "query": "Is our top video ad fatigued or is it just auction CPM inflation?", "should_trigger": false },
{ "query": "Write 10 ad copy variants from this value proposition", "should_trigger": false },
{ "query": "Draft a creative brief for a UGC video campaign", "should_trigger": false },
{ "query": "Design a creative test plan with per-cell budgets and kill rules", "should_trigger": false },
{ "query": "Mine our search terms report for negative keywords", "should_trigger": false },
{ "query": "Which negative match type should I use at the account level?", "should_trigger": false },
{ "query": "Are we pacing to spend our $60K monthly ad budget?", "should_trigger": false },
{ "query": "Should we switch from manual CPC to target ROAS bidding?", "should_trigger": false },
{ "query": "Split our $200K quarterly budget across Google and Meta", "should_trigger": false },
{ "query": "Set a maximum allowable CAC policy for the whole org", "should_trigger": false },
{ "query": "Compute our blended CAC and judge whether it's healthy", "should_trigger": false },
{ "query": "Which paid channels fit a $5K/month budget for a B2B tool?", "should_trigger": false },
{ "query": "Should we consolidate our 40 ad sets without resetting learning?", "should_trigger": false },
{ "query": "When has our winning campaign earned a budget increase?", "should_trigger": false },
{ "query": "Audit the landing page our ads send traffic to", "should_trigger": false },
{ "query": "Score these five video hooks and tell me which deserve budget", "should_trigger": false },
{ "query": "Build a swipe file of competitor ads organized by hook type", "should_trigger": false },
{ "query": "Is a carousel the wrong format for our lead-gen objective?", "should_trigger": false },
{ "query": "Plan paid promotion of our CEO's LinkedIn posts", "should_trigger": false },
{ "query": "Write UGC scripts with three hook variants for our skincare brand", "should_trigger": false },
{ "query": "Adapt our ad copy for AI assistant answer placements", "should_trigger": false },
{ "query": "How do I move from PPC specialist to paid media lead?", "should_trigger": false },
{ "query": "Design the interview loop for hiring a senior media buyer", "should_trigger": false },
{ "query": "Which paid media newsletters and podcasts should I follow?", "should_trigger": false },
{ "query": "What's my seed round pitch missing? We're raising $2M", "should_trigger": false },
{ "query": "Generate seed data for our staging database", "should_trigger": false },
{ "query": "Best practice for setting a random seed in reproducible ML experiments?", "should_trigger": false },
{ "query": "Run an RFM segmentation for our email retention campaigns", "should_trigger": false },
{ "query": "Which customers should get our loyalty program invite based on LTV?", "should_trigger": false },
{ "query": "Build lookalike models in Python to predict churn", "should_trigger": false },
{ "query": "Find companies similar to our best customers for cold outbound", "should_trigger": false },
{ "query": "Clean our email list to improve deliverability", "should_trigger": false },
{ "query": "Set up a suppression list so we stop emailing unsubscribers", "should_trigger": false },
{ "query": "Do we need a cookie consent banner for our EU site?", "should_trigger": false },
{ "query": "Segment our audience personas for content marketing", "should_trigger": false },
{ "query": "How do I grow my podcast audience?", "should_trigger": false },
{ "query": "Improve our SEO by analyzing search audience intent", "should_trigger": false },
{ "query": "Which influencers have audiences similar to our customer base?", "should_trigger": false },
{ "query": "A/B test our landing page headline", "should_trigger": false },
{ "query": "How should we structure our email welcome sequence?", "should_trigger": false },
{ "query": "Identify our ICP from our customer data", "should_trigger": false },
{ "query": "Export our CRM contacts to a CSV for the sales team", "should_trigger": false },
{ "query": "What's a good survey panel size for statistical significance?", "should_trigger": false },
{ "query": "Estimate the sample size for our conversion A/B test", "should_trigger": false },
{ "query": "Set up reverse ETL from our warehouse to Salesforce", "should_trigger": false },
{ "query": "GDPR data-deletion workflow for our SaaS product", "should_trigger": false },
{ "query": "Hash user passwords - bcrypt or argon2?", "should_trigger": false },
{ "query": "Match and dedupe duplicate contacts inside our CRM", "should_trigger": false },
{ "query": "Which audience should our webinar invite email go to?", "should_trigger": false },
{ "query": "Plan the audience Q&A segment for our conference talk", "should_trigger": false },
{ "query": "Increase our Instagram followers organically", "should_trigger": false },
{ "query": "Winsorize the outliers in our revenue forecast model", "should_trigger": false },
{ "query": "Choose a CDP vendor for our martech stack", "should_trigger": false },
{ "query": "Our display retargeting frequency is annoying users, cap it", "should_trigger": false }
]
}
references/examples.md›
# Worked Seed Specifications
Two worked examples: a B2C value-based seed that clears its floor directly, and a B2B closed-won seed that fails the floor and takes the fallback ladder. Numbers are illustrative of the method. The platform floors and match ranges they use are the documented ones.
## Example 1 - B2C ecommerce, value-based seed (clears the floor)
Context from the interview:
- Skincare brand, ~48,000 lifetime customers in the warehouse.
- Identifiers: email, phone, name, postal on ~70% of rows.
- Value column: margin, computed by finance.
- EEA customers present. Consent for ad-platform matching is captured at checkout, and opt-outs sync nightly to the warehouse.
Destination: a value-based audience on a large social platform. Goal: new-customer CPA under $34, with acquired-customer 90-day LTV at or above the account average of $61.
Selection:
- RFM Champions + Loyal, 180-day window (buyers are pixel-tracked), margin as the value column.
- Negative-selection pass applied: 2,100 rows removed (refunders, chargebacks, discount-code-only buyers, staff domains, two wholesale accounts). Result: 7,400 rows.
- Expected match: 60% - platform's good-list band is 50-80%, list is fresh and multi-identifier, so mid-band.
- Effective seed: 7,400 x 0.60 = ~4,440 matched, clears the 100 floor and sits inside the 1,000-5,000 recommended band.
- US-only campaign, so no country split needed.
```
SEED SPECIFICATION - champions-loyal-margin v1, 2026-08-20
platform : social (value-based customer list) | audience type: value-based
definition : identity-resolved customers, RFM Champions+Loyal, last purchase <= 180 days
value column : contribution margin (USD, finance-owned, refunds excluded) | mapping confirmed: at upload
row count : 7,400 | identifiers/row: email, phone (E.164 digits), first/last name, zip
expected match : 60% (platform good-list band 50-80%; fresh multi-identifier list)
effective seed : 7,400 x 0.60 = ~4,440 matched vs floor 100 -> clears (within 1,000-5,000 recommended)
fallback used : none
exclusions : refunders, chargebacks, discount-only, employees, wholesale stripped from seed;
all-customers suppression audience excluded at delivery level
consent basis : checkout consent for ad matching; opt-outs/deletions enforced in warehouse nightly (2026-08-19)
refresh : bi-weekly via warehouse sync | owner: lifecycle marketing lead
measurement : new-customer CPA <= $34; cohort LTV at 90/180/365d by seed version vs $61 account avg
re-check : 2026-11-20 (90-day cohort checkpoint; learning phase respected before any early read)
```
## Example 2 - B2B closed-won seed (fails the floor, takes the fallback ladder)
Context: infrastructure-software vendor, 410 closed-won customers, average contract $38K, CRM has work emails only for most contacts. Destination: predictive audience on the professional network (300 matched floor). An earlier attempt uploaded the raw list and the audience never served - the team read 410 rows as a 410 seed.
First computation: 410 contacts, work emails on the professional network match at ~30-40%, so take 35%. Effective seed: 410 x 0.35 = ~144 matched. **Fails the 300 floor.**
The spec does not get loosened to "all opportunities including lost" - that is trading selection quality, which the ladder forbids as a first move.
Fallback ladder. Rung 1 leads here despite being the most expensive rung in effort: a vendor, a DPA, and counsel sign-off, so a week or more of coordination.
This is the case where the cheap rungs are structurally empty, which the ladder's efficiency default calls out. The list is already all-time closed-won, so there is no recency left to widen - rung 2 costs an afternoon and returns nothing. The only adjacent segments are lost deals or MQLs, which the selection-quality rule forbids as a first move.
On a B2C list with 48,000 candidates the order would have been the opposite: rung 2 first, and procurement probably never started.
1. **Enrich identifiers** (rung 1): identity enrichment appends personal emails and phones, and fills missing contacts on won accounts via title/seniority/function expansion - 410 accounts become 1,240 contacts at a projected 70% match. Effective seed: ~868 matched. Clears 300. **Stop here** - rung 1 resolved it. Rungs 2-4 stay unused.
2. (Not needed, and empty here) Widen recency one notch - the seed is already all-time, so there are no older closed-won deals to add.
3. (Not needed) Stack an adjacent segment - late-stage open opportunities of the same segment.
4. (Not needed) Switch to a non-purchase source - SQL-only demo requesters, never all-MQL.
```
SEED SPECIFICATION - closed-won-enriched v2, 2026-08-22
platform : professional network (contact list -> predictive audience) | audience type: predictive
definition : contacts on closed-won accounts (all-time), expanded by title/seniority/function,
enriched with personal identifiers
value column : none (flat seed; ACV band is homogeneous by selection)
row count : 1,240 | identifiers/row: work email, personal email (enriched), phone
expected match : 70% (enrichment vendor's projected band; native work-email match was ~35%)
effective seed : 1,240 x 0.70 = ~868 matched vs floor 300 -> clears
fallback used : ladder rung 1 (identifier enrichment + account-to-contact expansion);
selection quality unchanged - still closed-won only
exclusions : churned-with-refund accounts, partner/reseller accounts, employees;
current-customer domains excluded at delivery level
consent basis : legitimate-interest assessment on file per counsel; enrichment vendor DPA signed;
opt-outs enforced in CRM export job - counsel sign-off dated 2026-08-15
refresh : monthly manual export until warehouse sync ships | owner: revops manager
measurement : cost per SQL from the audience; pipeline created; cohort read only at sales-cycle
maturity (~2 quarters) - early reads limited to match quality and engagement
re-check : 2026-09-22 (30-day match/engagement check; cohort verdict deferred to maturity)
```
The contrast to carry over: Example 1's constraint was selection discipline (48,000 candidates, pick ~7,400). Example 2's constraint was the floor (410 candidates, reach 300 matched without diluting to lost deals or MQLs). Same arithmetic, opposite binding constraint - which is the B2C/B2B split in practice.
references/platform-floors.md›
# Platform Floors, Match Rates, and Refresh Behavior
Figures verified against official documentation or attributed practitioner sources as of **August 2026**. These numbers move - Google cut Customer Match from 1,000 to 100 in 2025, and Demand Gen reach tiers convert to signals around March 2026 - so verify against the platform's live documentation before finalizing any spec. Where a platform's own pages disagree, both figures are shown. **Plan against the stricter one**.
## Floors and recommended sizes
All floors apply to the **matched** count, not uploaded rows.
| Platform | Matched floor | Recommended seed | Expansion tiers | Notes |
| ------------------------------------------------- | --------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Meta | 100 matched people (customer list); 100 unique conversions (conversion-based) | 1,000-5,000 (Meta developer docs and Business Help Center); 200+ converters for conversion-based seeds | 1%-10% of a chosen country (API allows up to 20% in 1-point steps) | Country-scoped by construction. "Advantage+ lookalike" replaced lookalike expansion; inclusions behave as suggestions. The widely repeated 1,000-50,000 sweet spot is practitioner guidance, not Meta's figure. |
| Google Customer Match | 100 members (cut from 1,000 in 2025) | 1,000+ for a usable optimization signal (practitioner) | Similar Audiences REMOVED Aug 1, 2023 (new segments stopped generating May 2023) | SHA-256 hashing. Members not added/refreshed within 540 days drop out of eligibility (official). |
| Google Demand Gen lookalike | 1,000 active matched (API docs); Help Center also cites 100 - **take 1,000** | - | Narrow 2.5% / Balanced 5% / Broad 10% of target location | Country-scoped; refreshes every 1-2 days; reach tiers become signals ~March 15, 2026. |
| LinkedIn Matched Audiences / Predictive Audiences | 300 matched members to serve | Contact list: upload 10,000+ emails; company list: 1,000+ companies (LinkedIn recommendation) | Lookalikes SUNSET Feb 29, 2024 -> Predictive Audiences (300 min, max 30 per ad account) | Contact list 300-300,000 rows, 20MB cap. Company lists dedupe on match, so padding rows is harmless. |
| TikTok | 1,000 for custom audience/lookalike source (Help Center); FAQ mentions 100 after mapping - **take 1,000** | 10,000 (practitioner) | Narrow / Balanced / Broad (~1% / 5% / 10%) | 24-48h processing; one Help article says a lookalike cannot be built directly from an email list. |
| Pinterest | 100 (Actalike source and customer list) | 1,000+ (practitioner) | Actalike 1%-10% (API caps at 10%; countries US/CA/GB) | Auto-refreshes as the source changes. |
| Reddit | 1,000 matched (25 via LiveRamp); max 1 million | - | Native reach expansion via automated targeting | Native CSV match ~2% (pseudonymous accounts); enrichment vendors claim 50-70%. Accepts hashed email and hashed mobile-ad-ID. |
| X | 100 matched to target (lowered from 500); audiences under 500 not shown in the interface | - | Follower look-alikes; custom-audience lookalikes | Match rate defined as 90-day active users / provided (official). |
| Amazon DSP / AMC | AMC lookalike seed 500 user IDs; rule-based audiences 2,000; DSP activation floors ~1,000 | 10,000+ for stable retargeting; 100,000+ for prospecting reach (practitioner) | AMC/DSP lookalike expansion | Matched counts below ~1,000 are rounded or hidden for privacy (documented by sync-tool vendors). |
| Snapchat | 1,000 to run an ad | - | Similarity / Balance / Reach | Lookalike updates as the seed updates. |
| Microsoft Ads | ~1,000 historical Customer Match minimum; Similar Audiences deprecated | - | - | Consistent with the industry-wide similar-audience deprecation trend. |
## Deprecations - stale-playbook markers
- **Google Similar Audiences**: removed from all ad groups and campaigns starting August 1, 2023. Any plan citing them is outdated. First-party lists now act as signals for optimized targeting and Demand Gen lookalikes.
- **LinkedIn Lookalike Audiences**: sunset February 29, 2024, replaced by Predictive Audiences (300-member minimum, max 30 per ad account).
- **Meta detailed-targeting exclusions**: removed March 31, 2025. Only custom-audience exclusions under Audience Controls remain hard suppression rules. Custom-audience inclusions became suggestions.
## Expected match rates
Use these to compute the effective seed. Adjust down for old lists and single-identifier rows.
| Platform | Typical match rate | Source type |
| ------------------------------------------------ | -------------------------------------------------------------------------------- | ----------------------------------------- |
| Google Customer Match | 29-62% "most advertisers"; 70-90% clean lists; phone-only 30-50% | Official (Google Ads Help) + practitioner |
| Meta | 50-80% good lists; below 40% signals a poor/outdated list; 30-60% commonly cited | Practitioner, attributed to Meta guidance |
| LinkedIn | ~30-40% on work-email contact lists; higher for company lists | Practitioner + LinkedIn Help |
| TikTok / consumer platforms with B2B work emails | Under ~5% on raw CRM work-email exports; ~40-85% after identity enrichment | Practitioner (enrichment vendors) |
| Reddit | ~2% native CSV; 50-70% after enrichment | Practitioner |
| Amazon DSP | 40-70% hashed email | Practitioner |
**What lifts match rate**, ordered by efficiency - `normalize formats > join identifiers already held > drop rows past 18 months > buy identity enrichment`:
- **Normalize the formats already in hand**: minutes, a transform in the export query, reversible, no approvals, and it lifts every row at once. The most common single cause of a rate below the platform's documented band, so always first.
- Emails lowercased and trimmed.
- Phone as digits with country code, no plus sign, no leading zero (E.164-style, e.g. 14155552671).
- Never pre-hash a UI upload, since platforms hash on ingest and a salted hash never matches.
- **Join identifiers the business already holds** - phone, name, location sitting in another table. Multiple identifiers per row beating email alone is the documented multiplier, and here it costs about an hour of joining and no procurement. Email stays the anchor identifier.
- **Drop rows past 18 months** - minutes, but a trade rather than a gain: the rate rises partly because the denominator shrinks. Never run it on a list already near its floor, and never let it flatter a spec whose effective seed then fails.
- **Buy identity enrichment** - a week or more of coordination (vendor selection, security review, sign-off, match test), plus a DPA, a lawful-basis re-check that covers enrichment and not only upload, and the hardest reversal of the four once contracted. Lowest ratio in general; unbeatable in exactly one case, a B2B work-email list on a consumer platform, where the three levers above cannot move a ~5% native match anywhere near a floor.
Meta exposes an Event Match Quality (EMQ) score with a practitioner target of 6.0+. Re-rank against the user: an enrichment contract already signed moves the last lever to the top, since both its setup and its DPA are already paid for. A warehouse holding no second identifier anywhere removes the second lever entirely.
## Refresh behavior
| Mechanism | Auto-refreshes? | Guidance |
| --------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
| Pixel/engagement-sourced audiences | Yes (Meta lookalikes roughly every 3-7 days; Demand Gen every 1-2 days; Pinterest as the source changes) | Keep the source rule healthy; nothing manual needed |
| Static CSV uploads (LinkedIn lists, manual Meta customer lists) | No | Practitioner cadence: weekly for active customer-list campaigns, bi-weekly to monthly for lookalike seeds, weekly for suppression lists |
| Google Customer Match membership | No - members expire | Members not added or refreshed within 540 days become ineligible (official) |
| Warehouse/CDP automated sync (reverse ETL) | Yes - adds and removes continuously | The dominant modern pattern; manual CSV is a stale-list risk with a named-owner maintenance burden |
Choosing between the two mechanisms the advertiser controls: `automated warehouse sync > static CSV` on efficiency, even though the CSV wins on effort - roughly a week of engineering once, against near-zero to start and then a standing job forever. The sync also wins on compliance cost, which is the argument that usually settles it: it propagates opt-outs and deletion requests continuously, while a manual CSV re-uploads yesterday's opt-outs every cycle and puts the lawful basis back in question each time. The ordering flips only where no engineering capacity exists at all, and there the honest answer is a CSV on a shorter cadence, not a sync that never ships.
Data older than 12-18 months produces both lower match rates and worse performance - decay both ways.
references/seed-selection.md›
# Seed Selection: Cohorts, Windows, Exclusions, Value, and Fallback Sources
## RFM selection - the default
RFM (Recency, Frequency, Monetary) is the dominant named framework for picking seed rows. The operationalized version (Klaviyo) splits customers into six mutually exclusive cohorts: Champions, Loyal, Recent, Needs attention, At risk, Inactive - with the Monetary score defined on historic customer lifetime value rather than AOV. The canonical seed export is **Champions + Loyal combined**.
Predictive layers on top where available: predicted CLV (needs roughly 500 orders of history to be usable), churn risk, next expected purchase date.
Top-slice heuristics when RFM tooling is absent (attributed practitioner rules). Efficiency ordering, best ratio first - `pattern filter > ticket/frequency slice > revenue slice`:
- **Pattern filter** - 2+ orders, AOV above a threshold, first order more than 90 days ago. Built specifically to strip recent one-time buyers. Best ratio on the board: one WHERE clause, no percentile computation, no coordination, and it encodes the negative-selection intent directly instead of hoping the percentile does.
- **Top 20-25% by average ticket or purchase frequency.** About an hour of SQL - a per-customer aggregate plus a percentile. Approximates worth better than spend does.
- **Top 10% by revenue over the past 12 months**, defined in the BI layer or SQL and validated with the finance/data team before export. Same query effort as the slice above, but the validation adds days of coordination, and revenue drifts toward the tenure/low-margin failure the value-based rules below warn about. Lowest ratio - use it when finance already maintains the definition, which flips its effort to near-zero.
All three assume no RFM tooling and no analyst on call. If the user has either, re-rank: with a CLV model already in place, none of these compete with predicted CLV.
Expressed as a query, a monetary seed is a per-customer aggregate with a threshold:
- Group orders by resolved customer identity.
- Sum order value.
- Keep rows above the cutoff.
- Always select against identity-resolved records, not raw source rows, so one person never appears as three.
Why quality-first: practitioner consensus, stated verbatim by Stackmatix, is that "a 150-person high-LTV customer list outperforms a 15,000-person newsletter subscriber list because it gives Meta a tighter signal cluster." Grow With Sakib reports 500 paying customers beating 5,000 newsletter subscribers "in every meaningful test". Directional practitioner claims, not audited studies - but consistent across sources, and aligned with Meta's own homogeneity-over-size guidance.
## Recency windows
- **30-90 days**: the sourced starting point for customer-list seeds.
- **180 days**: acceptable for pixel/CAPI-tracked buyers (also the platform-side maximum lookback for website custom audiences on Meta).
- **12-18 months**: the decay cliff - contact data older than this produces materially lower match rates and worse performance. Drop or archive these rows.
- Widening the window is fallback rung 2 - never part of the initial selection, and only reached once the effective seed has actually failed a floor. It may be run before rung 1 when no enrichment vendor is contracted, since it costs an afternoon. That is a cheaper order of attempts, not a licence to open with a wider window.
## The negative-selection pass
Run before every export. Strip from the seed:
- Refunders and chargeback customers
- Serial returners
- Discount-only buyers (never purchased at full price)
- Employees (by company domain or an internal list)
- Wholesale / reseller accounts
- Anyone on the do-not-contact, opt-out, or deletion list (enforced in the warehouse, before any sync)
Keep existing customers **in** the seed - they are the signal - but suppress them from the acquisition campaign at delivery level with a separate exclusion audience. Excluding past customers from the seed itself removes exactly the people the model should learn from.
## Value-based seeds
A value-based seed adds a per-row value the platform weights the model on. Rules:
- **Feed margin or predicted LTV, never cumulative lifetime revenue** - cumulative revenue tracks tenure, and high-revenue low-margin customers poison the model.
- Positive values only. Zero and negative values break or distort the build.
- Values must not be all identical - identical values carry no ranking signal.
- One currency per list. Platforms normalize scale internally but not mixed currencies.
- Cap or winsorize whales so a handful of outliers doesn't dominate the weighting.
- Exclude refunded amounts from the value computation.
- **Confirm the value column mapped at upload.** An unmapped column is a silent failure: the platform builds a standard lookalike and nothing warns you. If the value-based option is absent at audience-build time, the mapping failed (documented by MHI Growth Engine and LeadEnforce as the most common value-based mistakes, alongside confusing order count with value).
- Meta's stated requirements: minimum 100 matched, aim 1,000+. Value-based audiences need a separate terms-of-service acceptance.
## Non-purchase seed sources - ranked by efficiency
When purchase data is too thin to clear a floor at acceptable quality, pick from these eight. Rows are ordered by **efficiency** - signal quality bought per unit of setup effort - not by the sourced practitioner value ranking, which stays visible as `V1`-`V8` because it is the sourced artifact and effort runs almost exactly backwards to it.
| Source | Value | Effort | Consent exposure |
| ------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| Purchasers | V1 - best per-row signal by a wide margin | An hour **when the customer export already exists**, the usual case; a data project when it doesn't - warehouse export, identity-resolution join, normalization | Full: lawful basis, suppression, platform list terms |
| High-intent page viewers (pricing, demo) | V3 | Near-zero if the pixel already fires; about an hour with one judgment call to agree which pages count as intent | None new when built platform-native; full gate if exported to a list |
| Lead-form submitters - submitters, never openers | V4 | Near-zero on a platform-native form; about an hour on an owned form the pixel tracks | None new when platform-native |
| Checkout / add-to-cart initiators | V2 | An hour if commerce events are instrumented and trusted; a tracking project rather than a seed choice if they aren't | None new when platform-native |
| Email clickers - never openers; Apple Mail Privacy Protection (2021) inflated opens into an unusable signal | V5 | About an hour: ESP export, click segment, dedupe against the warehouse | Re-enters the full gate, and ESP consent scope rarely covers ad-platform upload - check it, never assume it |
| Video viewers, by completion: 95% > 75% > 50% > 25% > 3-second. 75% is the named inflection point - below it intent drops off sharply | V6 | Minutes, inside the ad account, off signal the platform already holds | None - nothing leaves the platform |
| Page/profile engagers | V7 | Minutes, same | None |
| All site visitors | V8 | Minutes | None |
Purchasers keep the top slot in the common case: a large value gap for one query against a table someone maintains anyway. Where the export genuinely doesn't exist, high-intent page viewers and lead-form submitters are the best ratio on the board - near-zero setup, no upload, no new consent exposure, and still in the top half on value. Below lead-form submitters, per-row quality falls off faster than effort does, so stop letting effort pull the choice down: a cheap seed that teaches the model the wrong customer costs more than the hour it saved.
Platform-native engagement audiences are also durable against mobile-OS tracking loss, because the signal originates inside the platform, with retention windows up to 365 days. In B2B, an SQL-only seed beats an all-MQL seed at every rung.
Re-rank before recommending:
- A warehouse with identity resolution already built, or a data team that owns the customer export on a schedule, drops the purchaser row's effort to near-zero and widens its lead.
- No pixel and no commerce events removes the middle rows from the board.
- An unresolved consent basis leaves only the zero-exposure rows available at any effort.
- A hard launch date promotes every minutes-not-weeks row.
- A compounding-asset mandate promotes the purchaser row even when its export has to be built.
## B2B specifics
- **Seed from closed-won CRM opportunities**, not MQL exports. Practitioner note: most B2B teams find 300-1,000 closed-won records sufficient for initial testing.
- **The identity problem**: work emails match under ~5% natively on consumer platforms, and personal emails match poorly on LinkedIn. Identity enrichment (resolving work identities to personal identifiers, or the reverse) is the documented workaround - it turns rung 1 of the fallback ladder into the standard B2B move.
- **Documented case**: a tight 200-account, ~600-contact ABM list matched ~180 people natively on LinkedIn and never activated below the 300 floor. Enrichment to 70-90% match made a 400-500 contact list viable.
- **Account-level vs contact-level**: expand a company list into contacts using job title + seniority + function filters before upload. LinkedIn: company list 1,000+ companies recommended, contact list roughly 10,000+ emails uploaded to reliably clear 300 matched.
- **Segment mixed account lists** into homogeneous bands (enterprise / mid-market / SMB) - left as one audience, delivery over-serves the largest companies in the list.
- Firmographic/technographic sources (industry, headcount, revenue, installed technology) can define the account list that becomes the seed, but the seed itself should still be selected on won revenue, not on fitting the ICP on paper.
## Splitting the seed
Each split becomes its own seed and its own lookalike. Split in this order - `country > product line > high-AOV vs. all > language/region`:
- **Country** - not really a choice: lookalikes are country-scoped, so a single global seed doesn't produce a worse lookalike so much as an averaged one that serves no market. Do it before considering any other split.
- **Product line or SKU category with different buyers** - the split that changes what the model learns most, so the highest-value optional one.
- **High-AOV vs. all customers** - a test rather than a decision. Its value is one comparison, so run it only once the two above are settled, then keep the winner.
- **Language/region** - usually correlates with country and adds the least new information.
Effort doesn't rank this list: every split is one more WHERE clause. What ranks it is the standing cost, which is identical per split and permanent - one more audience to refresh on cadence, monitor for overlap, and hold above its floor forever, plus one more chance of silently falling under it.
Two splits is fine. Six is a part-time job nobody was assigned. Cap the count at what the named refresh owner can maintain, and check each split's own effective seed against the floor **before** committing, not after.
The test for splitting: would a single model trained on the combined list be learning two different customers? If yes, split. Homogeneity matters more than size, so splitting trades size for homogeneity, usually the right trade above the floor.
Re-rank against the user: an automated warehouse sync collapses every split's standing cost to near-zero and makes more splits affordable, while a manual CSV process makes the third split cost as much attention as the first two combined.
SKILL.md›
---
name: lookalike-audience-seeds
description: "Select and size the seed customer list behind a lookalike, similar, or value-based audience - which customers to upload, how many, and whether the matched count (not the row count) clears the platform floor - with RFM and value-based selection, a negative-selection pass, a privacy and consent gate before any customer list upload, and a fallback ladder when the floor cannot be cleared. Use whenever the user mentions a lookalike or similar audience, a seed list, a customer list upload, match rate, an audience that is too small, or says their lookalike isn't working - even if they never say 'seed'. Covers B2B and B2C. Do NOT use to design the whole targeting tier structure - use mbfinotti/advertising-skills@ad-audience-targeting instead."
license: MIT
metadata:
author: Maya-Beth Finotti
version: "1.3.6"
---
# Lookalike Seeds
Select, size, and validate the seed customer list that a lookalike, similar, or value-based audience will be built from. Three principles govern every step below:
- **Match count, not row count** - every uploaded list shrinks by its match rate before it reaches the platform's minimum. A 40% match on 5,000 rows is a 2,000-person seed, and the platform floor applies to that matched number - upload 2-3x the target.
- **Homogeneity and value concentration beat size** - Meta's own guidance holds that a seed's homogeneity affects audience effectiveness more than its size. Practitioner consensus (Stackmatix, Grow With Sakib - directional, not audited) is that a few hundred high-value customers outperform thousands of undifferentiated subscribers.
- **Inclusion is a suggestion, exclusion is a rule** - on Meta (Advantage+) and Google (Demand Gen, where lookalike reach tiers now function as signals rather than hard segments), the delivery algorithm may override an inclusion audience. Only exclusion/suppression audiences remain hard rules.
The seed's job is feeding the algorithm the cleanest signal, not fencing an audience - which raises the quality bar rather than removing it.
This skill ends at a specified, validated seed definition plus its export/handoff spec. It never walks through creating or uploading the audience inside any ad platform's interface.
Where each adjacent concern lives:
- Overall targeting plan (interest, behavioural, lookalike, and retargeting tier layering) - seed mechanics live here, tier architecture lives in `mbfinotti/advertising-skills@ad-audience-targeting`.
- Retargeting sequence design, audience windows per funnel stage, and frequency caps - `mbfinotti/advertising-skills@retargeting-funnel`.
- B2B buying-committee and persona mapping - `mbfinotti/advertising-skills@ad-buyer-group-mapper`.
- Root-causing broad account underperformance - `mbfinotti/advertising-skills@ad-account-diagnostic`.
- Pixel/CAPI instrumentation feeding engagement-based seed sources - `mbfinotti/advertising-skills@ad-conversion-tracking`.
- Reconciling conflicting performance reads on a seed across attribution sources - `mbfinotti/advertising-skills@ad-attribution-gap`.
## Interview
Ask before selecting anything. One question per message; offer multiple-choice answers where possible; skip anything already answered or visible in supplied data.
- B2B, B2C/ecommerce, or both?
- Which platform(s) will consume the seed? (Meta / Google / LinkedIn / TikTok / other - floors and match mechanics differ)
- What customer data exists, and where does it live? (CRM / data warehouse / email platform / pixel-and-platform events only / almost none)
- Roughly how many total customers, and which identifiers exist per row? (email / phone / name / postal address / mobile ad ID - more identifier types per row lift match rate)
- Is there a usable value column - and what does it actually measure: cumulative lifetime revenue, margin, predicted LTV, or just order count? Provenance matters more than presence.
- How recent is the data - what share of the list purchased or converted within the last 90 and 180 days?
- Does the list contain EEA/UK residents or US state-privacy-covered consumers? Is there a documented consent or lawful basis for ad-platform upload, and are opt-outs, deletions, and do-not-contact records enforced before export?
- How will the list reach the platform - manual CSV or automated warehouse/CDP sync? Who owns the refresh?
- What acquisition outcome is the lookalike meant to drive - target new-customer CPA/CAC, and any LTV or margin expectation for acquired customers?
- What does the account currently average on those same two numbers - acquired-customer CPA, and acquired-customer cohort LTV at 90 days? The pass threshold grades the seed against this baseline, so an unknown baseline means the seed cannot be graded at all: say so and set the first cohort checkpoint as the baseline-building run.
- By what date must the audience be live and spending, and is that date hard? (Seed sources span minutes to a week of coordination, and the enrichment rung runs through procurement - the orderings in steps 3 and 6 cannot be picked without this.)
- Do you want the audience that can launch soonest, or a compounding asset - an automated warehouse sync, enriched identifiers, a reusable cohort definition - that keeps paying across every future audience?
- Effort ceiling: how many hours of data-team time are available, who owns the refresh afterwards, and can you actually get privacy sign-off and a vendor DPA? (A "no" deletes rungs rather than delaying them.)
## Workflow
Every ranking in this skill - steps 3, 6 and 9 here, and the ones in the references - is a default, not a law. Each assumes the common case: no enrichment vendor under contract, no in-house identity graph, no analyst free this week. Who executes it changes the ordering as much as what it is.
Re-rank first against the interview's deadline, durability, and effort-ceiling answers, naming which answer moved which rung:
- A hard near-term date promotes the platform-native engagement seeds.
- A compounding-asset mandate promotes the warehouse sync and the enrichment rung despite their setup.
- No route to privacy sign-off deletes every list-based rung outright.
Then re-rank against what this user actually owns - any of these collapses a rung's effort to near-zero and moves it up:
- A signed enrichment DPA.
- Identity resolution already built in the warehouse.
- An analyst who can write the cohort SQL today.
- A data team that owns the export job anyway.
1. Run the Interview; collect every answer before selecting a single row.
2. Run the privacy and consent gate - a hard stop, before any selection work. If any item below is unconfirmed, stop and say exactly what is missing. This skill is not legal advice: lawful basis varies by jurisdiction, so route open questions to the user's counsel.
- Documented lawful basis or consent for uploading this list to an ad platform. Under GDPR, the Bavarian DPA position upheld by the Higher Administrative Court Munich (2018, Ref. 5 CS 18.1157) holds that customer-list custom audiences require prior consent, and hashing does not change that: SHA-256 output is pseudonymous, not anonymous, so consent obligations survive it.
- For EEA use of Google Customer Match, consent signals passed per Consent Mode v2 with both ConsentStatus fields GRANTED (mandatory since March 2024).
- CCPA/CPRA "sharing" opt-outs and Global Privacy Control signals (honored in twelve US states) propagated before export.
- No seed built on sensitive categories: health, sexual orientation, religion, race, political affiliation, criminal history.
- Platform customer-list terms accepted for the specific ad account (value-based audiences need a separate terms acceptance on Meta).
- Suppression, opt-outs, and deletions enforced once, in the warehouse, before any sync - not per-platform afterwards.
3. Choose the seed source: purchase data first, always. If purchase data is too thin, pick among the eight sourced alternatives - but never down the value ranking alone, because effort and compliance cost both run backwards to it. The axes disagree, so read all four:
- value: `purchasers > checkout/add-to-cart initiators > high-intent page viewers > lead-form submitters > email clickers > video viewers > page engagers > all visitors`
- effort: `video viewers == page engagers == all visitors > high-intent page viewers == lead-form submitters > checkout initiators > email clickers > purchasers` - the platform-native engagement audiences build inside the ad account in minutes off a rule the pixel already fires, and skip the customer-list upload and its step-2 consent gate entirely. A purchaser or CRM seed needs a warehouse export, an identity-resolution join, and the full gate. Both ties are real equalities: the first three are one blanket rule over an event already firing, and the next two are that same single rule pointed at a named URL set or a named form - minutes either way, by the same person.
- compliance cost: `purchasers == checkout initiators == email clickers > lead-form submitters > high-intent page viewers == video viewers == page engagers == all visitors` - anything reaching the platform as a customer list runs the full step-2 gate (lawful basis, Consent Mode v2 in the EEA, opt-out propagation, a separate terms acceptance for value-based audiences), while a platform-native engagement audience never leaves the ad account and triggers none of it. Both ties are exact rather than approximate: the three list-based sources run the identical gate item for item, and the four native ones run none of it - this gate has no partial version to separate them by. The axis can decide the pick on its own: where consent for upload is unconfirmed, every list-based rung is unavailable, not merely expensive.
- efficiency: `purchasers > high-intent page viewers > lead-form submitters > checkout initiators > email clickers > video viewers (75%+) > page engagers > all visitors`. Purchasers lead **when the customer export already exists**, the common case - a large value gap for one query. Where it doesn't, high-intent page viewers and lead-form submitters are the best ratio on the board: near-zero setup, no upload, no new exposure, still top-half on value. Those two are ordered rather than tied - their setup is the same single rule, so the value axis breaks it - except in B2B lead gen, where the form submission _is_ the conversion and submitters move ahead. Never let effort pull the choice below them: past that point per-row quality falls off faster than effort does. Per-source breakdown and the conditions that move this order: [references/seed-selection.md](references/seed-selection.md).
What this efficiency order starves is the customer-identity asset itself: the purchaser or CRM seed, and the identity enrichment that makes it usable (fallback rung 1). It tops the value axis and sits last on both effort and compliance cost, so wherever the export does not already exist it loses every round to a pixel rule that ships in minutes - and the account still has no export the next time it asks.
Promote it against the ratio when:
- The interview answered "compounding asset" rather than "live soonest".
- The seed is B2B and the native match rate cannot reach any floor without enrichment.
- This same account has now lost the round twice.
Delete it instead - by name, out of the menu - where there is no route to privacy sign-off or no vendor DPA available. A list-based source with no lawful basis is not a slow option, it is not an option, and one left ranked at the bottom reappears as scope after the audience is already live.
Use email clickers, never openers - Apple Mail Privacy Protection (2021) inflated opens into noise. In B2B, seed from closed-won CRM opportunities, not all-MQL exports. SQL-only beats all-MQL.
4. Select the rows. [references/seed-selection.md](references/seed-selection.md) has the full cohort definitions, window guidance, and exclusion list.
- Default selection: RFM Champions + Loyal cohorts, or the top 10-25% by predicted LTV or margin - never by cumulative lifetime revenue, which tracks tenure, not worth.
- Recency window: 30-90 days as the starting point, up to 180 for pixel/CAPI-tracked buyers. Contact data older than 12-18 months materially degrades both match rate and performance.
- Negative-selection pass (most teams forget this): strip refunders, chargebacks, serial returners, discount-only buyers, employees, and wholesale accounts. Keep existing customers in the seed, but suppress them from the acquisition campaign at delivery level.
5. Compute the effective seed: `effective seed = uploaded rows x expected match rate`. Pull the expected match rate from the platform's documented range in [references/platform-floors.md](references/platform-floors.md), adjusted for the list's identifier density and age - for example:
- Google Customer Match: 29-62% for most advertisers (per Google's own docs).
- Meta: 50-80% on good lists.
- LinkedIn: ~30-40% on work-email contact lists.
- Reddit: ~2% native CSV.
Compare the result to the platform's matched-count floor and size the upload at 2-3x the target. If your harness can read the customer export, compute the counts directly. Otherwise, emit the exact filter definitions (source table, filter conditions, recency bounds, value threshold) for the user's data team to run, and wait for the counts before proceeding.
6. If the effective seed cannot clear the floor, walk the fallback ladder - never by loosening selection quality first. Rung numbers are stable IDs used by the spec block and the worked examples; each rung carries its own value, effort, and compliance cost:
- **Rung 1, enrich identifiers** - identity resolution appending personal emails/phones. Value: the only rung that adds matched people without changing what the model learns, and on a B2B work-email list against a consumer platform the only rung that reaches any floor at all from a ~5% native match. Effort: the outlier the ladder's numbering hides - vendor selection, security review, budget sign-off, then a match test, so a week or more of coordination before a single row moves. Compliance: a new processor in the chain - DPA, counsel sign-off, and a lawful-basis re-check covering enrichment and not just upload; the hardest rung to reverse once contracted.
- **Rung 2, widen the recency window one notch** (90 to 180, 180 to 365). Value: moderate, paid for in signal freshness. Effort: near-zero - a query edit, minutes, fully reversible. Compliance: none new.
- **Rung 3, stack an adjacent segment of the same homogeneity** - add Loyal to Champions, or a second product line's buyers if behaviour is genuinely similar. Value: moderate, and homogeneity is what's at risk. Effort: about an hour, plus the judgment call that is the actual work. Compliance: none new.
- **Rung 4, switch to the next non-purchase seed source down the step-3 ranking.** Value: lowest - the only rung that changes what the model learns. Effort: about an hour if the event source is already instrumented. Compliance: **lower**, not higher - a platform-native source needs no upload and no consent gate.
Efficiency by default: `rung 1 > rung 2 > rung 3 > rung 4`. Stop at the first rung that clears, and record it in the spec.
One documented flip: `rung 2 > rung 1`, when no enrichment vendor is contracted **and** there is recency left to widen - an afternoon may end the problem before procurement starts, and rung 1 stays available if it doesn't. The flip does not apply on the standard B2B case, where the cheap rungs are structurally empty: an all-time closed-won list has no recency left, no adjacent segment that isn't lost deals, and a native match too low for rungs 2-4 to reach any floor.
Rung 1 is also the rung to delete rather than demote. No procurement route, no DPA, or no owner for a vendor budget takes it off the ladder entirely - say it is deleted and run rungs 2-4. A rung nobody can execute, left sitting at the top of an efficiency order, is how the spec ends up assuming it.
On the standard B2B case that leaves no ladder at all: report that as the finding instead of filling it with a looser seed.
7. If the seed is value-based:
- Feed margin or predicted LTV.
- Positive values only.
- Values must not be all identical.
- One currency per list.
- Cap or winsorize whale outliers so a few large accounts don't skew the model.
- Confirm the value column is mapped at upload. An unmapped value column makes the platform silently build a standard, not value-based, lookalike - if the value-based option doesn't appear at audience-build time, the column didn't map.
8. Specify identifier preparation:
- Multiple identifiers per row beat email alone.
- Phone in E.164-style digits with country code, no plus sign, no leading zero.
- Emails lowercased and trimmed.
- Never pre-hash a file destined for a UI upload - the platform hashes on ingest, and a salted or double hash will never match.
On Meta, target an Event Match Quality score of 6.0+ as the practitioner benchmark.
9. Decide splits, refresh, and overlap.
**Splits.** Split in this order - `country > product line > high-AOV vs. all` - and stop when the named refresh owner runs out of capacity. Effort per split is identical (one WHERE clause), but each adds a permanent one: an audience to refresh, monitor for overlap, and keep above its floor forever.
Country is not a choice: Meta lookalikes are country-scoped by construction, so multi-country coverage needs one seed-and-lookalike pair per country. Check every split's own effective seed against the floor before committing to it.
**Refresh.** Set the refresh cadence with a named owner. Static CSV uploads do not auto-refresh (LinkedIn lists, manual Meta customer lists), so schedule weekly-to-monthly refreshes or move to an automated warehouse sync.
Between those two, `automated sync > manual CSV` on efficiency and on compliance cost, despite losing on effort. A week of engineering once retires the whole decay failure mode and propagates opt-outs and deletions continuously. The CSV is near-zero to start but becomes a standing job forever, re-uploading yesterday's opt-outs every cycle.
Recommend the CSV only as a bridge, with a named owner and a date the sync lands. Google Customer Match drops members not refreshed within 540 days.
**Overlap.** Check overlap between this seed's audience and existing ones. Practitioner concern threshold is roughly 20-30%. Consolidate or exclude above it.
10. Fill the Seed Specification block (shape below; worked B2C and B2B versions in [references/examples.md](references/examples.md)) and hand the seed's tier placement to `mbfinotti/advertising-skills@ad-audience-targeting`.
11. Attach the measurement plan and a re-check date one cohort window out. If your harness has persistent memory, memorize the seed version, its definition, counts, and the re-check date so the next run compares against history instead of re-deriving it.
## The Seed Specification
Deliver one block per seed:
```
SEED SPECIFICATION - <seed name/version>, <date>
platform : <destination platform(s)> | audience type: <lookalike | similar | value-based | predictive>
definition : <source records> filtered by <selection rule> within <recency window>
value column : <none | column, what it measures, currency> | mapping confirmed: <yes | at upload>
row count : <n rows> | identifiers/row: <email, phone, ...>
expected match : <x%> (basis: <platform range + list quality adjustment>)
effective seed : <rows x match = n matched> vs floor <platform floor> -> <clears | fails>
fallback used : <none | ladder rung + what changed>
exclusions : <negative-selection list applied; suppression audiences delivered separately>
consent basis : <basis + suppression enforced in warehouse date | BLOCKED - missing item>
refresh : <cadence + owner + mechanism (manual CSV | automated sync)>
measurement : <new-customer CPA target, cohort LTV checkpoints by seed version>
re-check : <date - first cohort checkpoint>
```
## Evidence Gate
Refuse to emit a Seed Specification built on any of the following; state what is missing instead:
- **Unknown match rate basis** - no platform range consulted and no prior upload to calibrate against. Estimate only from the documented ranges plus list age/identifier density, and label it an estimate.
- **Unconfirmed consent basis** - the privacy gate (workflow step 2) did not pass. No selection quality justifies an upload the user has no lawful basis to make.
- **Value column of unknown provenance** - "we have an LTV field" is not evidence. If nobody can say whether it is revenue, margin, or a model output, the seed is not value-based until someone can.
Never invent a customer list, a segment definition, or a threshold the user's data does not support - request the missing evidence rather than substituting a generic list.
## Seed Diagnosis Ladder
When an existing lookalike underperforms and the seed is suspected, use this table as a symptom lookup, not a ranked menu: each symptom has one cause and one action, so ranking them against each other would be false precision.
What does need an order is which to fix when several symptoms fire at once. Fix anything that stops the audience serving or matching before anything about its quality - a floor failure, a broken mapping, or a match rate under the platform's band makes every quality read downstream of it meaningless. Then fix overlap and refresh, then the value/expansion causes.
| Symptom | Likely cause | Action |
| ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------- |
| Match rate far below the platform's documented range | Stale list, single identifier per row, or unnormalized formats | Normalize (E.164, lowercase email), add identifier columns, drop rows older than 18 months, then enrich |
| List uploaded but audience won't build or serve | Matched count under the floor - row count misread as seed size | Recompute effective seed; run the fallback ladder |
| Value-based option absent at audience build | Value column not mapped during upload | Re-upload with mapping confirmed; the platform silently built a standard lookalike |
| CPA fine, 90-day cohort LTV of acquired customers weak | Seed scaled low-margin converters (the documented Churney pitfall) | Re-seed on margin or predicted LTV; retire the seed version |
| Lookalike performs near random / no better than broad | Expansion percentage too wide, or heterogeneous seed mixing unlike buyers | Tighten the percentage; split seed by product line, segment, or country |
| Plan cites Google Similar Audiences or LinkedIn Lookalikes | Stale playbook - Similar Audiences removed August 2023; LinkedIn Lookalikes sunset February 2024 | Rebuild on Customer Match as signal / LinkedIn Predictive Audiences (300 matched min) |
| Two seeded audiences cannibalizing each other | Audience overlap above the ~20-30% concern band | Consolidate seeds or add mutual exclusions |
| Performance decays over weeks with no change made | Static CSV never refreshed; members aging out | Set cadence and owner; automate the sync; respect the 540-day Customer Match cutoff |
## Platform Floors
Floors apply to the **matched** count. Compact view - full table with recommended sizes, expansion tiers, documented internal inconsistencies, deprecations, match-rate and refresh behavior in [references/platform-floors.md](references/platform-floors.md). Platform floors move over time: if your harness can browse the web, verify against the platform's live documentation before finalizing a spec; otherwise instruct the user to verify.
| Platform | Matched floor | Recommended seed |
| ----------------------------- | ------------------------------------------------------------------------- | ----------------------------------------------------------------- |
| Meta | 100 | 1,000-5,000 (Meta docs) |
| Google Customer Match | 100 (cut from 1,000 in 2025) | 1,000+ (practitioner) |
| Google Demand Gen lookalike | 1,000 active matched (API docs; Help Center says 100 - take the stricter) | - |
| LinkedIn (Matched/Predictive) | 300 matched | contact list: upload 10,000+ rows; company list: 1,000+ companies |
| TikTok | 1,000 (Help Center; FAQ says 100 - take the stricter) | 10,000 (practitioner) |
| Pinterest | 100 | 1,000+ (practitioner) |
| Reddit | 1,000 matched | - |
| X | 100 matched | - |
| Amazon AMC lookalike | 500 user IDs | 10,000+ (practitioner) |
| Snapchat | 1,000 to serve | - |
## B2B vs B2C
**Identical in both:** hashing and normalization mechanics (SHA-256, E.164, lowercase email), match-rate arithmetic, the effective-seed computation, negative selection, warehouse-first suppression, and the measurement windows - they run on the same ad systems. What differs is list size and identity.
**B2B** lives with the small-list problem: 400 high-ACV customers cannot feed a consumer-platform lookalike natively, because work emails match poorly on Meta and TikTok while personal emails match poorly on LinkedIn. Identity enrichment before upload is the documented workaround, and the reason it is fallback rung 1.
B2B seeds are account-level or contact-level: expand a company list into contacts via title/seniority/function before upload. On LinkedIn, a company list wants 1,000+ companies, while a contact list needs roughly 10,000+ emails uploaded to reliably clear the 300-match floor. Seed from closed-won, not from MQLs.
The long sales cycle means the conversion read on a B2B seed lags by weeks: judge early on match quality and engagement, judge the seed itself only at cohort maturity.
**B2C** gets the opposite trade: lists large enough that selection discipline, not floor-clearing, is the binding constraint.
## Measuring Whether This Worked
Match rate is a data-quality leading indicator only - per Google's own wording, a high match rate suggests correct data but does not guarantee list performance. Never grade a seed on it.
Judge a seed version on new-customer CPA/CAC, new-to-brand rate, AOV of acquired customers, and the skill's own KPI: **cohort LTV at 90/180/365 days by seed version**. The named trap (Churney): low-CPA lookalikes that underperform on LTV at maturity because the seed scaled low-margin converters. The fix is re-seeding on margin or predicted LTV and retiring the version.
Compare seed variants with a user-level split test or a geo holdout, never a naive ad-set comparison - audience overlap contaminates it. Between the two valid methods:
- `split test > geo holdout` on effort - a platform setting versus withholding a market's traffic for the whole test window, which needs sign-off from whoever owns that market's number.
- `geo holdout > split test` on evidence strength - it survives cross-device identity loss and measures incremental sales, not platform-attributed ones.
Default to the split test where the platform offers a true user-level one. Escalate to the geo holdout when the seed decision is expensive to reverse - a re-platformed value column, a retired seed version, a country split - or when the split test's own read is what's in dispute.
Respect the learning phase before reading anything: roughly 50 optimization events per ad set per week (widely attributed to Meta's guidance) or 2-4 weeks at low volume.
Pass threshold: the seed version's 90-day acquired-customer cohort LTV meets or beats the account's average acquired-customer LTV at equal-or-better CPA. Below that, iterate: re-seed on value, tighten recency, or split. Re-check one cohort window later - a seed is a versioned artifact, not a one-time upload.
## Common Failure Modes
Deliberately unranked: these are trap/fix pairs, not competing options for one goal, and every fix here is cheap enough that ordering them by efficiency would imply a triage choice that does not exist - do all of them.
| Trap | Why it burns | Fix |
| ------------------------------------------------------- | ---------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------- |
| Sizing to the row count | The floor applies to matched people; a 5,000-row list at 40% match is a 2,000 seed | Compute effective seed; upload 2-3x the target |
| Padding the seed with everyone to "hit the minimum" | Dilutes the signal cluster; homogeneity outweighs size | Run the fallback ladder instead - quality is never the first thing traded |
| Feeding cumulative lifetime revenue as the value column | It tracks tenure, and high-revenue low-margin customers poison the model | Use margin or predicted LTV |
| Skipping the negative-selection pass | Refunders, discount-only buyers, employees, wholesale teach the model the wrong customer | Strip them before export, every time |
| Pre-hashing a file for UI upload | The platform hashes on ingest; a salted hash matches nothing | Upload normalized plaintext via the UI; hash only in API/warehouse syncs that require it |
| Uploading before consent is confirmed | Hashed lists are still personal data under GDPR; enforcement is real | Privacy gate first; suppression enforced once, in the warehouse |
| Grading the seed on match rate | Match rate measures data hygiene, not audience quality | Grade on cohort LTV by seed version |
| Judging a seed inside the learning phase | Early CPA is noise until optimization volume accumulates | Wait for learning exit or 2-4 weeks; then read |
| One global seed for a multi-country lookalike | Lookalikes are country-scoped; the model averages across markets | One seed-lookalike pair per country |
| Treating the upload as done forever | Static lists decay; members age out (540-day rule on Customer Match) | Refresh cadence with an owner, or automated sync |
| Mixing work and personal identity per platform | Work emails barely match on consumer platforms; personal emails barely match on LinkedIn | Match identifier type to platform; enrich when they diverge |
| Copying 2022-era playbooks | Similar Audiences and LinkedIn Lookalikes no longer exist; inclusions became suggestions | Verify features against live docs; treat the seed as signal, exclusions as the only hard rule |
## Reference
- Read [references/platform-floors.md](references/platform-floors.md) when sizing against a floor - full per-platform table with dates, inconsistencies, deprecations, match-rate ranges, and refresh behavior.
- Read [references/seed-selection.md](references/seed-selection.md) when selecting rows - RFM cohorts, recency windows, the negative-selection list, value-based rules, and the non-purchase source ranking.
- Read [references/examples.md](references/examples.md) when writing the spec - a worked B2C value-based seed and a worked B2B closed-won seed that fails the floor and takes the fallback ladder.