research

Sequence Performance

Email campaign/sequence performance review composite. Pulls campaign data (sends, opens, replies, bounces), reads actual email copy and subject lines, analyzes reply content (objections, positive interest, questions), and produces a diagnostic report covering quantitative metrics, copy quality, lead quality, and actionable recommendations. Tool-agnostic — works with Smartlead (MCP), Instantly, Outreach, Lemlist, Apollo, or CSV data.

Gooseby Athina AI
Install
Terminal
npx gooseworks install --all

# then, in Claude Code, Cursor, or Codex:
/gooseworks use the sequence-performance skill
About This Skill

Sequence Performance

Goes beyond vanity metrics. Most campaign reports tell you open rate and reply rate. This skill reads the actual emails you sent, reads every reply you received, classifies the responses, evaluates your copy, evaluates your lead quality, and tells you specifically what's working, what's not, and what to do about it.

Three layers of analysis:

  1. Quantitative: The numbers — sends, opens, replies, bounces, conversions, by touch and by variant
  2. Qualitative (Copy): Are the subject lines, email bodies, CTAs, and personalization actually good?
  3. Qualitative (Replies): What are people actually saying? What objections keep coming up?

When to Use

Use this skill when:

  • User says "how's my campaign doing", "sequence performance", "campaign review", "email analytics"
  • User says "analyze my outreach", "why isn't my campaign working", "review my email results"
  • A campaign has been running for 7+ days and has meaningful data

Phase 0: Intake

Outreach Tool

  1. What outreach tool do you use? (Smartlead / Instantly / Outreach.io / Lemlist / Apollo / Other)
  2. How do we access campaign data? (MCP tools / API / CSV export / paste metrics)

Campaign Selection

  1. Which campaign? (name or ID)
  2. Date range? (or "all data")

Your Company Context (for copy evaluation)

  1. What does your company do? (one-liner)
  2. Who is your ICP? (titles, industries, company size)
  3. What problem do you solve?
  4. What's your CTA goal? (book meeting, get reply, drive to page)

Benchmark Context

  1. Is this cold outreach or warm/nurture?
  2. What segment are you selling to? (SMB, mid-market, enterprise)

Step 1: Pull Campaign Data

Pull three categories of data from the user's outreach tool:

A) Campaign Metrics

Data PointWhat We Need
Total emails sentBy touch (Touch 1, Touch 2, Touch 3, etc.)
Total unique recipientsDeduplicated count
OpensBy touch, unique opens vs. total opens
RepliesBy touch, total reply count
BouncesHard bounces + soft bounces
UnsubscribesCount
ClicksIf link tracking is on
Positive repliesIf categorized in the tool
Meetings bookedIf tracked

How to pull by tool:

ToolMethod
Smartlead (MCP)mcp__smartlead__get_campaign_stats, mcp__smartlead__get_campaign_sequence_analytics, mcp__smartlead__get_campaign_variant_statistics
Instantly / Outreach / Lemlist / ApolloAsk user for CSV export or paste metrics
OtherUser provides CSV with columns: email, status, opened, replied, bounced

B) Email Copy (Sequence Content)

Pull the actual templates for every touch:

ToolMethod
Smartlead (MCP)mcp__smartlead__get_campaign_sequences
OthersUser pastes the copy or provides CSV export

C) Reply Content

Pull the actual text of every reply:

ToolMethod
Smartlead (MCP)mcp__smartlead__get_campaign_leads_history, mcp__smartlead__fetch_master_inbox_replies
OthersUser provides reply dump or CSV export

Human Checkpoint

Campaign: [name]
Status: [active/paused/completed]
Sent: X emails to Y recipients
Replies: Z (full text pulled for analysis)
Touches: N touches, M variants
 
Data looks complete? (Y/n)

Step 2: Quantitative Analysis

Benchmarks

MetricCold (SMB)Cold (Mid-Market)Cold (Enterprise)Warm/Nurture
Open rate40-60%30-50%25-40%50-70%
Reply rate3-8%2-5%1-3%10-20%
Positive reply rate1-3%0.5-2%0.3-1%5-10%
Bounce rate<3%<3%<2%<1%
Unsubscribe rate<1%<1%<0.5%<0.5%

Calculate

Overall metrics: open rate, reply rate, positive reply rate, bounce rate, unsubscribe rate, deliverability rate. Compare each to the benchmark.

Per-touch breakdown:

  • Touch-level open/reply rates
  • Marginal reply rate (replies from THIS touch / people who received this touch but hadn't replied yet)
  • Touch contribution (what % of total replies came from each touch)

Variant analysis (if A/B testing):

  • Open rate and reply rate per variant
  • Statistical confidence: <50 sends = "insufficient data", 50-100 = "directional", 100-250 = "likely winner", 250+ = "statistically significant"
  • Winner recommendation: scale, keep testing, or kill

Step 3: Reply Analysis

Read every reply, classify it, and extract patterns.

Reply Categories

CategoryDefinition
Positive interestWants to learn more, open to a conversation
Meeting requestExplicitly asks to meet or provides availability
Warm / CuriousInterested but non-committal, asks questions
Objection — TimingNot now, but potentially later
Objection — BudgetCan't afford or not a priority
Objection — CompetitorAlready using a competing solution
Objection — RelevanceDoesn't see the fit
Objection — AuthorityNot the right person
Not interestedFlat no
Auto-reply / OOOAutomated response
ReferralRedirects to someone else
QuestionAsks about product/offering

Objection Patterns

  • Which objection appears most? (reveals systemic issues)
  • Do objections cluster at Touch 1 (bad targeting) vs. Touch 3 (fatigue)?
  • Which are handleable (timing, authority) vs. terminal (relevance)?
  • What exact language do people use?

Positive Signal Patterns

  • Which touch/variant generated positive replies?
  • What do positive responders have in common? (title, industry, company size)
  • What questions do warm leads ask? (reveals what's missing from the email)

Reply Quality Score

ScoreCriteria
Strong>50% positive/warm. Objections are handleable.
Mixed30-50% positive. Mix of handleable and terminal.
Weak<30% positive. Dominated by "not interested" and "not relevant."
ToxicHigh unsubscribe + angry replies. Something is fundamentally wrong.

Step 4: Copy Quality Assessment

Evaluate the actual email copy against best practices and reply data.

Subject Lines

CriterionRed Flags
Length>60 chars gets truncated on mobile
SpecificityGeneric "Quick question" or "Checking in"
Spam triggers"Free", "Limited time", ALL CAPS
Open rate correlationLow open rate = subject line problem

Email Body

CriterionRed Flags
Hook (first line)"I'm reaching out because..." or "We are a company that..."
LengthOver 150 words
Value prop clarityJargon, vague language, buzzwords
Proof pointsNo proof = no credibility
PersonalizationOnly {first_name} merge field
CTAMultiple CTAs, high-friction asks, or no CTA
Filler language"Hope this finds you well", "just checking in"
Sequence progressionTouch 2 is just a "bump" of Touch 1

Grades

Grade each touch A through F on: hook quality, value prop clarity, proof usage, personalization level, CTA quality.

Step 5: Lead Quality Assessment

Evaluate whether we're sending to the right people.

Targeting Check

  • Do lead titles match ICP buyer/champion/user personas?
  • Are leads in target industries?
  • Right seniority level for the ask?
  • Company size in target range?

Signal Quality (from replies)

PatternWhat It Tells You
High "not relevant" repliesSending to people who don't have the problem
High "wrong person" repliesRight companies, wrong roles
High "already have a solution"Right problem, late to the party
High "timing" objectionsRight people, right problem, wrong moment — not a targeting issue
Low reply + high open ratePeople open but don't find it relevant — copy/targeting mismatch
High bounce rateList quality issue — bad emails, old data

Step 6: Generate Report

Report Structure

# Sequence Performance Review: [Campaign Name]
**Period:** [date range] | **Status:** [active/paused/completed]
 
---
 
## Executive Summary
 
**Overall verdict:** [One sentence]
 
| Dimension | Grade | Assessment |
|-----------|-------|-----------|
| Metrics | [A-F] | [one-liner] |
| Copy Quality | [A-F] | [one-liner] |
| Lead Quality | [A-F] | [one-liner] |
| Reply Quality | [Strong/Mixed/Weak/Toxic] | [one-liner] |
 
### What's Working (Double Down)
- [Specific thing with data]
 
### What's Not Working (Fix or Kill)
- [Specific thing with data]
 
### Top 3 Actions
1. [Highest-impact action]
2. [Second]
3. [Third]
 
---
 
## Detailed Metrics
 
### Overall Performance
| Metric | Actual | Benchmark | Status |
|--------|--------|-----------|--------|
| Open rate | X% | Y% | [above/below] |
| Reply rate | X% | Y% | [above/below] |
| Bounce rate | X% | <3% | [flag] |
| ... | ... | ... | ... |
 
### Performance by Touch
| Touch | Sent | Open Rate | Reply Rate | Marginal Reply Rate | % of Total Replies |
|-------|------|-----------|------------|--------------------|--------------------|
| 1 | X | Y% | Z% | Z% | W% |
 
### Variant Performance (if A/B testing)
| Touch | Variant | Subject | Sent | Open Rate | Reply Rate | Confidence | Action |
|-------|---------|---------|------|-----------|------------|------------|--------|
 
---
 
## Reply Deep Dive
 
### Reply Classification
| Category | Count | % of Replies |
|----------|-------|-------------|
 
### Top Objections
| Objection | Count | Handleable? | Suggested Response |
|-----------|-------|------------|-------------------|
 
### Notable Replies
[5-10 most instructive replies with quotes]
 
---
 
## Copy Assessment
[Subject line verdicts, body grades, sequence architecture assessment]
 
---
 
## Lead Quality
[Targeting assessment, actual vs intended ICP]
 
---
 
## Recommendations (Prioritized)
 
### High Priority (Do This Week)
1. **[Action]** — [data point] → [expected impact]
 
### Medium Priority (Do This Month)
2. **[Action]** — [data point] → [expected impact]
 
### Kill List
- [Anything that should be stopped]

Recommendation Logic

FindingRecommendation
Open rate below benchmarkSubject line rewrite — suggest 3 alternatives
Reply rate below + open rate fineBody copy issue — focus on hook, proof, CTA
Both below benchmarkFull sequence rewrite
High "not relevant" objectionsTargeting issue — tighten ICP filters
High "wrong person" referralsTitle targeting issue — shift to referred titles
High "already have solution"Add competitive differentiation to copy
High "timing" objectionsNot a problem — set up 90-day re-engagement
One variant clearly winningScale winner, test new idea in losing slot
Touch 2/3 near-zero marginal repliesCut sequence short or rewrite with new angles
High bounce rateList hygiene — verify emails, check data source
Deliverability <95%Infrastructure — check SPF/DKIM/DMARC, reduce volume

Human Checkpoint

Present the executive summary, then ask:

Full detailed report available. Want to see the full breakdown, or act on a specific recommendation?

Adapting to Data Availability

Missing DataWhat Gets SkippedStill Useful?
Reply textReply classification + objection patternsPartially — metrics + copy still run
Variant dataVariant analysisYes — single-variant analysis still runs
Lead demographicsTargeting assessmentYes — infers from reply patterns
Open trackingOpen rate analysisPartially — reply rate + copy still run

Minimum viable data: Emails sent + reply count + email copy text.

Cost

Free. Pure reasoning + data from user's outreach tool.

Tips

  • Run at Day 7 and Day 14. Day 7 catches deliverability and subject line problems. Day 14 gives enough replies for objection analysis.
  • Reply analysis is where the gold is. Metrics tell you WHAT. Replies tell you WHY.
  • High open + low reply = copy problem. The subject gets them to open but the email doesn't deliver.
  • Low open + decent reply rate = subject line problem. The email works, people just aren't seeing it.
  • "Not relevant" is the most important objection. If >20% say "this isn't for me," it's targeting, not copy.
  • Don't kill a variant too early. Need 100+ sends per variant for directional data.
  • Touch 2/3 should contribute 30-40% of replies. If Touch 1 is 90%+, your follow-ups aren't adding value.

What's included

·
User says "how's my campaign doing", "sequence performance", "campaign review", "email analytics"
·
User says "analyze my outreach", "why isn't my campaign working", "review my email results"
·
A campaign has been running for 7+ days and has meaningful data
·
Touch-level open/reply rates
·
Marginal reply rate (replies from THIS touch / people who received this touch but hadn't replied yet)
You Might Also Like

Render VO Anchored Motion Listicle

Assemble an expert/educator motion-graphic LISTICLE video ad from a config — a spoken authoritative voiceover carries a numbered listicle while N web-animated hyperframe beats (HTML plus the Web Animations API, one branded design system of alternating tiles, big hero numerals, and glass-pill callouts) are rendered frame-by-frame via Playwright and anchored to the VO's word-level timestamps, periodic color-graded B-roll windows give visual breath, and captions burn ONLY inside those B-roll windows (2-word chunks, ASS Format header carrying a Name field so none drop) with the VO mixed under a low music bed. This is the FREE deterministic assembly stage (Playwright beat render plus ffmpeg concat plus window-masked caption burn plus VO-and-music mix plus final composite) — the VO, the music bed, and the stock B-roll come from create-vo-elevenlabs, create-music-elevenlabs, and media-proxy. Use for the vo-anchored-motion-listicle format.

Render Stopmotion Hand Swatch Cycle

Assemble a stop-motion hand-swatch-cycle product-demo ad from a config — a sequence of still PLATES (one hand swiping a single-barrel cosmetic across a cream skin-patch, the barrel + swatch changing per plate while the hand, background, crop, and lighting stay locked) is PNG→mp4 loop-encoded at each plate's own stop-motion hold (fast motion frames 150–250ms, per-shade ~380ms, hero beats 1100–1800ms), concat-demuxed with HARD cuts into a silent master, closed on a Playwright HTML-rendered branded end card (serif tagline + sans subtitle + real logo SVG over a hero BG, never AI-rendered text), and muxed with a pre-sourced music track playing under the end card with a fade tail (no VO). This is the FREE deterministic assembly stage (loop-encode + concat-demux + end-card render + music mux); the master-anchor plate, shade plates, and end-card BG come from create-image-gpt-image-fal and the track from create-music-elevenlabs. Use for the stopmotion-hand-swatch-cycle format.

Render Split Screen Creator

Assemble a split-screen creator ad from a config — a two-zone vertical composite where a supplied AI-creator lip-sync take fills the BOTTOM ~48% while real 16:9 product/demo clips run uncropped in the TOP ~52%, each top clip contain-fit with a darkened blurred cover-scale fill of the same clip (never black bars), a 3px brand-color divider between the zones, the creator slice cover-fit per the per-scene VO timing, scenes hard-concatenated with the body audio being the concatenated creator VO slices, an end card held on the last sharp frame ~3s, then the ASSEMBLED cut transcribed with local Whisper (not the raw VO — concat drops inter-scene silence) and word-level captions burned in the chosen style. This is the FREE deterministic assembly + caption stage (two-zone composite + blurred fill + divider + hard-concat + end card + captions); the VO comes from create-vo-elevenlabs, the anchor from create-image-gpt-image-fal, and the whole-VO lip-sync from a paid VEED Fabric 1.0 take (a no-atom upstream input). Use for the split-screen-creator format.

Newsletter

Learn to build Growth systems with AI

2-3 compounding systems per week using Claude Code, OpenClaw, and more.