4/10/2026
How to A/B Test WhatsApp Sales Scripts: 50 Samples, Four Variables
Before you change a single line of your sales script, isolate the sentence you want to test and freeze everything else. If you skip that step, a bump in reply rate tells you nothing about what caused it. In short, how to A/B test WhatsApp sales scripts is about giving reps leverage, not replacing them.
Why gut-feel script changes make things worse
On Monday you swap in a new opener, send 30 messages, and get 8 replies. You conclude the new line works. But Monday at 10 a.m. is already a peak online window for your customers, and this batch happens to be high-intent leads you collected at a trade show last week. The reply bump may have nothing to do with your new sentence.
Three traps catch almost every team:
- Too few samples: You send 20 messages and call it. Two extra replies out of 20 pushes the reply rate from 10% to 20% — that is noise, not signal.
- Wrong attribution: A customer closes, and you credit the opener. What actually moved them was the phone call on day three. The script opens the door; the follow-up cadence walks through it.
- Survivorship bias: You remember the line that got an instant reply and forget the same line sent to 40 other people that went nowhere.
To control variables: use the same customer segment (all new inquiries), the same send window (weekday mornings at 10 a.m.), the same product, and change exactly one element — either the opener or the CTA, never both at once.
Define what you are actually testing
A script is not one block. Split it into four testable elements:
- Opening line: "Hi, are you there?" versus "I saw you checked out the XX product — which use case are you buying for?"
- Value proposition: price, lead time, or after-sales support.
- Call to action (CTA): "Reply if you're interested" versus "Send me your target quantity and I'll calculate the landed cost."
- Follow-up cadence: follow up after one day or three, and whether the follow-up carries new information.
Write one single hypothesis per element. For example: "Replacing 'Hi' with 'I saw you checked out the XX product' will raise the first-reply rate." Do not write "changing the opener and adding a quote will raise reply rate" — two variables move together and you learn nothing.
Lock your metrics in advance so you cannot cherry-pick afterward:
- First-reply rate: share of customers who reply within 24 hours of the message.
- Qualified-conversation rate: share who reply and ask a specific question (price, specs, lead time).
- Close rate: share of that batch who eventually order.
First-reply rate is the leading indicator; close rate is the ultimate one. With small samples, watch first-reply rate first. Once you have hundreds of sends, compare close rates.
Four steps to run one round inside WhatsApp
Step 1: Split. Use Excel or WhatsApp labels to randomly divide customers into groups A and B, at least 50 each. If you cannot send 50 in a day, accumulate over several days — but both groups must come from the same source. Do not put all old customers in A and all new inquiries in B.
Step 2: Send. Send both versions in the same time window to avoid time-of-day effects. For example, send version A to group one at 10:00 a.m. and version B to group two at 10:15 a.m. Use the sent/read status to track whether customers saw the message, and chat timestamps to record reply time so you can calculate the 24-hour first-reply rate.
Step 3: Log. Build a sheet where each row is one send, with columns: customer ID, group (A/B), script version, send time, replied within 24h, qualified conversation, closed. Log manually or with a tool — the key is not to miss rows.
Step 4: Decide. Group A: 50 sends, 10 replies, 20% first-reply rate. Group B: 50 sends, 17 replies, 34%. With 50 samples per group, a 14-point gap is worth adopting version B as the new standard and moving to the next round. If samples are under 30 — say A gets 3 replies from 12 and B gets 5 from 12 — do not conclude yet. Keep accumulating.
A full scenario: your current opener "Hi, we make XX" is version A. Version B becomes "I saw you checked out the XX product — which use case are you buying for?" Send 50 new inquiries to each group in the same time window. One week later: A gets 9 replies (18%), B gets 16 (32%). Same customer source, clearly higher first-reply rate for B. Adopt B as the new standard opener, then test the next element.
Iterate scripts with reply rate and close rate
Build a script iteration board and review it weekly. The table should include at least: script version, sends, first-reply rate, qualified-conversation rate, close rate, notes. Add new versions, keep old ones, so you can see how scripts evolve over time instead of starting from zero each week.
Optimize high-leverage elements first. Opening lines and CTAs move reply rates most directly — test those first. If reply rate is fine but close rate is low, the problem is in your value proposition or follow-up script. Shift testing to "how you quote" and "whether the follow-up carries new information."
Run at least 100 sends per version before evaluating. Switching after 20 sends lets noise completely mask the real difference.
Example: version B opener has a 32% first-reply rate but only a 3% close rate. Version A has 18% first-reply and 5% close. B attracts replies but the leads are less qualified. The next step is not reverting to A — keep B's opener and test a different CTA. Change "Reply if you're interested" to "Send me your target quantity and I'll calculate the landed cost" and see whether close rate improves.
How small teams without analysts can run this
Google Sheets is enough. Build a template with columns: send date, script version, customer group, sends, replies within 24h, qualified conversations, closes. Use formulas to auto-calculate first-reply rate (replies / sends) and close rate (closes / sends). Paste new data weekly and the board updates itself.
Hold a fixed 30-minute weekly review. Look at only two metrics: change in reply rate and change in close rate. Do not agonize over individual conversations or debate why one customer did not reply. Look only at the aggregate difference between the two groups. The meeting produces one output: which version to use next week and which new variable to test.
Write test results into an internal script library. Mark high-converting versions as "verified" so new hires can reuse them without re-testing from scratch. Organize the library by scenario: first touch on new inquiries, follow-up after quoting, reactivating silent customers, repeat orders from existing customers. Under each scenario, list verified versions and versions under test, and keep it current.
Manual sheets usually die from missed rows. A rep juggling 20 conversations sends a message and moves to the next chat; few go back to mark "this was version A, customer replied, counts as qualified." Incomplete data pushes testing back to guesswork. Some teams use a tool layered on top of WhatsApp Web to automatically archive replies and closes per script version, removing the manual tagging step. Sellenca's pricing is $19 per seat per month, or $190 per seat annually — compare that against the time cost of maintaining sheets by hand before deciding.
Breaking the deadlock when customers stop replying
First check whether you have tripped WhatsApp's risk controls. High-frequency sending of the same script can get throttled — messages show as sent but customers never receive them, or your reply rate suddenly drops from 30% to 5%. Lower your send frequency, reduce daily volume of the same script, or rephrase.
Test by segment. New customers, existing customers, and quoted-but-unclosed customers react completely differently to the same script. New customers may respond to "I saw you checked out the XX product," while existing ones respond better to "The model you asked about last time is back in stock." Test and log by customer type separately — do not use one script for everyone.
Introduce a language variable. When customers are not native English speakers, test translated versions. For Spanish-speaking customers, send a machine-translated version and a human-polished version separately; reply rates can differ noticeably. Multilingual testing adds variables, so fix the script content first and change only the language to isolate the difference.
From manual testing to systematic iteration
The biggest problem with manual logging is omission — incomplete data makes the test useless. A systematic approach solves three things: logging adds no extra burden on reps, data automatically attaches to customer profiles, and high-converting scripts feed the next round of hypotheses.
Sellenca mines Q&A and scripts automatically from real closed conversations — for example, which sentence actually led to an order when a customer asked about lead time. Those high-converting lines become ready-made version B hypotheses for your next test instead of guesses. It also generates a daily follow-up list telling you who to contact today and why, so data collection does not eat into selling time. You can see the specifics on the features page.
FAQ
How many samples do I need for a reliable A/B test?
At least 50 sends per version. If your daily volume is low, accumulate a week of data — but keep both groups from the same source and the same send window. Below 30 samples, treat results as directional only and do not make major script changes based on them. Above 100 sends, conclusions are relatively stable.
How do I avoid getting banned on WhatsApp while testing?
Control send frequency, avoid blasting identical content in a short window, and do not send spam. When using automation, make sure sending behavior resembles a human pace — for example, AI-suggested replies that require a rep to confirm before sending, rather than an unattended bot, carry lower risk of being flagged as spam.
Is reply rate enough, or do I need close rate too?
Both. Reply rate determines whether you get conversations; close rate determines whether you make money. Optimize in order: raise reply rate first, then improve close rate. If reply rate goes up but close rate does not move, your script is attracting unqualified customers — adjust the value proposition or CTA.
How does a small team with no historical data start its first A/B test?
Use your current script as version A and write an improved version based on common customer questions as version B. For example, if customers often ask "How long is lead time?", have version B proactively mention "Standard models ship in 7 days, custom in 15." Send 50 to each group and log replies manually. The goal of the first test is not to find the perfect script — it is to get the process running.
Script testing is not a one-off project. It is a weekly loop: form a hypothesis, split and send, log data, decide, update the script library. To see how replies and closes get logged automatically inside WhatsApp and how follow-up lists are generated, book a demo and run one A/B round with your own scripts.