Every engagement begins with an assessment and plan.

Arrange an assessment

Accelerator founders

How to test a growth channel in four to six weeks, and when to scale it

Most startups never finish testing a channel. Here is a four to six week test design, the arithmetic for sample size, a test log template and real results that show when to scale and when to cut.

Dineth Ratnayake

Founder of Codax · 21 January 2026 · 9 min read

Two people reviewing a test plan of sticky notes on a whiteboard

The short answer

To test a growth channel, change one variable against a control, fix the sample size in advance and run the test for four to six weeks against the same account list. Judge it on qualified conversations and pipeline, not opens, using a decision rule written before launch. Scale the winner by volume, then budget, then by stacking a second channel on the same accounts.

Key takeaways

  • Use the Bullseye framework: rank every channel, test the three most promising cheaply, then focus on the one that works.
  • A real test changes one variable against a control, runs a fixed sample for four to six weeks and follows a pre-written decision rule.
  • To read a 2x lift in reply rate, the control needs about 15 replies, which at a 5% reply rate is about 304 sends per version.
  • In two engagements, controlled test winners beat their controls by 2x to 4x, with the biggest gains from changing the sender and the audience.
  • Scale volume first, then budget, then stack channels, and cut a channel only after several full-sample tests fail.

Most startups do not fail to find a channel. They fail to finish testing one. A founder sends 60 emails, posts twice, runs ads for a week, sees nothing clear and moves on. Three months later, five channels have been "tried" and none has been tested.

This guide gives you a test design you can run in four to six weeks, the arithmetic for how many sends you need, a test log template, real results from controlled tests and the rules for when to scale and when to cut.

How do you test marketing channels as a startup?

Pick three promising channels, run a small, cheap, controlled test in each and put your effort behind the one that produces qualified conversations. Then keep testing inside that channel to improve it.

This is the Bullseye framework from Traction by Gabriel Weinberg and Justin Mares. The book's publisher describes it as a framework to test channels across nineteen traction channels, "from viral marketing to trade shows", and argues that startups should be putting half their resources into getting traction. In an excerpt published by Justin Mares, the five steps are brainstorm, rank, prioritise, test and focus.

  1. Brainstorm. Write down a realistic way you could use every channel, including the ones you think will not work.
  2. Rank. Sort channels into three columns: most promising now, could possibly work, long shots.
  3. Prioritise. Choose your inner circle, the three channels that seem most promising.
  4. Test. Run relatively cheap tests in each to learn what a customer costs, how many are available and whether they are the right customers.
  5. Focus. Put your resources behind the one channel that proved itself, and keep experimenting inside it.

The excerpt is explicit that tests should be small. Its example is running four ads rather than forty, with a rough answer from a few hundred dollars. The goal of the first round is direction, not scale.

For a B2B founder selling to companies, the inner circle is usually drawn from a short list: founder outreach on LinkedIn, email to a named account list, webinars or events, content in the founder's voice and retargeting to accounts that have already engaged. Pick the three where you already have an asset, such as a founder with a following, a partner programme or a customer who will speak.

Run the three tests in parallel, not in sequence. Three six-week tests run one after another take eighteen weeks. Three run side by side take six weeks, and you compare them against the same accounts in the same market conditions.

What does a good growth experiment framework look like?

A good growth experiment changes one variable against a control, runs a sample you fixed in advance for four to six weeks and is judged on qualified conversations and pipeline against a decision rule you wrote before it started. Anything less is an anecdote.

Design a channel test in seven steps

  1. Write the hypothesis

    One sentence: if we change X for this account list, Y will rise, because Z. Example: if the CEO sends instead of the company, replies will rise, because buyers answer peers.

  2. Change one variable

    Sender, segment, subject angle, offer, format or audience. Change two at once and you will not know which one worked.

  3. Keep a control

    Run the current version alongside the new one, at the same time, to the same kind of accounts. Split the list randomly.

  4. Fix the sample size

    Decide how many sends or impressions each version gets before you start. Use the arithmetic in the next section.

  5. Set a four to six week window

    Long enough for follow-ups to land and calls to be booked, short enough to make several decisions a quarter.

  6. Pick the success metric

    Qualified conversations booked and pipeline created. Use replies as the early read. Never judge on opens or clicks alone.

  7. Write the decision rule

    Before launch, write what result means scale, what means iterate and what means cut. Then follow it.

Keep the account list constant across the test. In the Codax method, every channel starts as a controlled four to six week experiment against an agreed account list, which is the only way to compare one channel with another fairly. See how we work for the full sequence.

Log every test in one table. The log is how you stop repeating tests you have already run and how you show investors that your growth is a system, not luck.

Test log template, filled with real test results

HypothesisVariableControlMetricResultDecision
CEO as sender lifts repliesSenderCompany senderReplies3x repliesScale
Leading with the CMIO books more meetingsFirst contactOther buyer titlesMeetings2.1x meetingsScale
Gated whitepaper drives qualified conversationsGatingUngated whitepaperQualified conversations3x qualified conversationsScale
Retargeting books cheaper meetingsAudienceCold titlesCost per meeting4x cheaper meetingsMove budget
Banking and SaaS reply more than insuranceSegmentInsuranceReplies2x repliesCut insurance
Your hypothesisOne changeCurrent versionCalls or pipelineAfter the full sampleScale, iterate or cut

The rows above come from the cybersecurity services and agentic AI healthcare engagements. The last row is the template for yours.

How many emails do you need to test cold outreach?

To read a doubling of reply rate, you need roughly 15 replies in the control version, which at a 5% reply rate is about 304 sends per version. To read smaller differences, you need far more. Most founders stop at a tenth of that.

Evan Miller gives a rule of thumb in How Not To Run an A/B Test: sample size n = 16 x variance / d squared, where d is the minimum effect you wish to detect. For a rate such as replies, the variance is p x (1 - p). His sample size calculator reports the answer per variation, so each version of your test needs that many sends.

For example, at a 5% reply rate, detecting a lift to 10% means d is 0.05. That gives 16 x 0.05 x 0.95 / 0.0025, which is 304 sends per version. Detecting a lift to 7.5% means d is 0.025, which gives 1,216 sends per version. Halving the effect you want to see quadruples the sample.

Sends needed per version to read a result

With a 2% reply rate in the control, reading a 2x lift takes about 784 sends per version, which is roughly 16 replies in the control.

Sends per version to read a 3x lift196
Sends per version to read a 2x lift784
Sends per version to read a 1.5x lift3,136

With a 5% reply rate in the control, reading a 2x lift takes about 304 sends per version, which is roughly 15 replies in the control.

Sends per version to read a 3x lift76
Sends per version to read a 2x lift304
Sends per version to read a 1.5x lift1,216

With an 8.5% reply rate in the control, reading a 2x lift takes about 172 sends per version, which is roughly 15 replies in the control.

Sends per version to read a 3x lift43
Sends per version to read a 2x lift172
Sends per version to read a 1.5x lift689

Worked example using Evan Miller's rule of thumb n = 16 x p(1 - p) / d squared, where p is the control reply rate and d is the lift you want to detect. 8.5% is the average outreach reply rate in Backlinko's study of 12 million emails.

The pattern in the tabs is the rule to remember. To read a 2x lift, the control needs about 15 replies whatever the reply rate. At 2%, that is about 784 sends per version. At 8.5%, the average outreach reply rate in Backlinko's analysis of 12 million outreach emails, it is about 172.

Miller's other warning matters just as much: do not peek and stop early. He writes that if you peek at an ongoing experiment ten times, what you think is 1% significance is actually just 5% significance. Fix the sample, let it run, then read it.

What should you measure instead of opens?

Measure qualified conversations and the pipeline they create. Use positive replies as the early signal, because they arrive inside the window, and confirm the winner on calls booked and pipeline.

Opens tell you a subject line was seen, not that a buyer cares. A version that wins on opens and loses on calls has lost. The Codax test phase for founders is judged on replies and calls booked, and every positive reply and signal reaches sales the same day so no result is lost to slow handling.

  • Early read (weeks one to three): positive reply rate per version, acceptance rate on LinkedIn, registrations for an event.
  • Decision read (weeks four to six): qualified conversations booked per 100 accounts reached, cost per qualified conversation.
  • Confirmation read (month two onward): pipeline created and stage progression from the winning version.

Paul Graham's advice in Startup = Growth is the reason to be strict. He writes that if you get growth, everything else tends to fall into place, so you can use growth like a compass. A compass only works if it points at the real number.

What did controlled channel tests actually show?

Across two engagements, every test winner beat its control by 2x to 4x on the metric that mattered. The biggest gains came from changing who speaks and who hears, not from changing words.

Test winners, cybersecurity services firm

Gated whitepaper over ungated (qualified conversations)3x
Founder-signed invitations (acceptances)3x
Founder ads over company ads (click-through)2.4x
Banking and SaaS over insurance (replies)2x
SOC operations webinar over compliance (registrations)2x

Codax, cybersecurity services firm, twelve-month engagement.

Test winners, agentic AI healthcare firm

Retargeting over cold titles (cost per meeting)4x cheaper
CEO as the sender (replies)3x
Health system speaker on webinars (registrations)2.6x
Leading with the CMIO (meetings)2.1x
Adding a user track (replies)2x

Codax, agentic AI healthcare firm, seven-month engagement.

Copy tests work the same way at accelerator stage. For an AI personalisation startup selling to ecommerce brands, outreach went from the founder's LinkedIn and email, testing angles on conversion, order value and repeat purchase. The team kept what booked calls and cut the rest. The engagement reached 17 brands in qualified pipeline worth $472K.

For a revenue cycle AI company acquiring RCM companies, the message about the owner's future (succession, staffing, margins and the cost of keeping up with technology) was tested with the founder until replies became calls. That work brought $8M ARR into design partnership from a $15M acquisition pipeline. The tests were judged on calls, not opens.

When should you scale a channel?

Scale a channel when a version has beaten its control on qualified conversations across a full sample, the cost per qualified conversation holds as you add volume and the accounts it reaches are the ones on your list. Then scale in three directions, in order.

  1. Volume. Send the winning version to more of the same kind of accounts. Watch cost per qualified conversation every month. If it rises sharply, you are reaching worse-fit accounts.
  2. Budget. Move spend from losing tests into the winner. A growth lead who owns the pipeline number should be able to move budget between channels within the quarter, not wait for next year's plan.
  3. Channel stacking. Add a second channel that reaches the same accounts. In the cybersecurity services engagement, 8 in 10 opportunities were touched by three or more channels, and accounts took an average of 5 touches before the first call.

Stacking is why a single-channel test undersells the winner. Outbound books calls faster when the buyer has already seen the founder's posts, a retargeting ad or a webinar invitation. Graham's advice on focus still applies: get one channel working before you stack the next.

Keep testing after you scale. Run one new challenger against the current winner every round, with the same sample discipline, so the channel keeps improving instead of decaying.

In the Codax method, the test phase runs from month two to month eight and scale starts from month five, with budget and volume moving to what converted. Pipeline is reviewed with leadership every month, account by account, so each scale decision rests on the pipeline rather than activity.

When should you cut a channel?

Cut a channel when two or three well-designed tests have each run their full sample and none has produced qualified conversations at a cost you can sustain. One failed test is a verdict on that version, not on the channel.

  • Cut the version when it loses to the control on the decision metric after the full sample.
  • Iterate the channel when replies arrive but calls do not. The offer or the follow-up is wrong, not the channel.
  • Cut the channel when the right accounts are reached, versions with real differences have been tested and qualified conversations still do not arrive.
  • Cut a segment, not the channel, when one segment wins and another loses, as banking and SaaS did against insurance.

Write the cut in the test log with the reason. Six months later, someone will suggest the channel again, and the log will save you a month.

How does a growth department run channel tests for a founder?

A growth department is one senior team that owns qualified pipeline end to end, from strategy to execution, under a single accountable lead. For founders, that team runs the test design, the lists, the sends and the log, while the founder writes and approves the copy and takes every call.

Copy is tested together every round, and what books calls gets more volume. A monthly pipeline review sets the next round. The founder keeps the voice and the closing, and the testing discipline runs whether or not the founder's week was full.

For the sequences to test, read the founder outbound playbook. For the mindset that keeps tests running long enough to read, see all the time in the world and no time at all. For the full arc from first test to repeatable pipeline, start with the YC growth playbook and the accelerator founders hub.

Questions and answers

How do you test marketing channels as a startup?

Use the Bullseye framework from Traction: brainstorm every channel, rank them, pick the three most promising, run cheap tests in each and focus on the one that works. Each test changes one variable against a control and runs a fixed sample for four to six weeks. Judge it on qualified conversations, not opens.

What is a growth experiment framework?

A growth experiment framework is a repeatable test design: a hypothesis, one variable, a control, a fixed sample, a set window, a success metric and a decision rule written before launch. Every result goes into a test log with the decision taken. The log turns individual tests into a system.

How many emails do you need to test cold outreach?

To read a doubling of reply rate, aim for about 15 replies in the control version. At a 5% reply rate, Evan Miller's rule of thumb gives about 304 sends per version. Reading a 1.5x lift at the same rate takes about 1,216 sends per version, so test bold changes on small lists.

How long should a channel test run?

Run a channel test for four to six weeks. That is long enough for follow-ups to land and calls to be booked, and short enough to make several decisions a quarter. Fix the sample in advance and do not stop early because a version looks ahead.

When should you scale a marketing channel?

Scale when a version has beaten its control on qualified conversations across a full sample and cost per qualified conversation holds as volume grows. Scale volume first, then move budget into the winner, then stack a second channel on the same accounts. Review pipeline monthly to confirm the scale is working.

Sources

  1. Strategize, Test, Measure: The Bullseye Framework (excerpt from Traction), Justin Mares, via Brian Balfour
  2. Traction by Gabriel Weinberg and Justin Mares, Penguin Books
  3. How Not To Run an A/B Test, Evan Miller
  4. Sample Size Calculator, Evan Miller
  5. Startup = Growth, Paul Graham
  6. We Analyzed 12 Million Outreach Emails. Here's What We Learned, Backlinko

Keep reading

More from the Newsroom

YC growth playbook
7 Jan 2026 · 10 min readAll the time in the world and no time at all: the founder mindset for speed and persistenceAct as if there is no time for everything you control. Persist as if there is all the time in the world for everything the buyer controls. Here are the rules, the evidence and the cadences.Read
YC growth playbook
22 Jul 2026 · 9 min readOutbound for founders: LinkedIn and email tips that book callsSmall lists, a warm-up before the first touch, three messages in the founder's own words and a waitlist that closes. Here is the sequence, step by step.Read
YC growth playbook
4 Nov 2025 · 9 min readThe YC growth playbook: from day one of the batch to a repeatable pipelineFour stages, from the first week of the batch to a pipeline a Series A investor trusts. The moves, the numbers and the mistakes at each one, with the founder kept as the voice throughout.Read

Start here

Every engagement begins with an assessment of what is already running

Findings shared in full, with a prioritised repair list, before anything is agreed.

What you receive

  • A written report of everything found
  • A prioritised repair list
  • A first read on the account list
  • A recommended plan across the five phases
Start with an assessment