Case Study

·

CPaaS

The Funnel Was Busy. The Pipeline Wasn’t.

How a CPaaS startup improved buyer fit without chasing more leads

A freelance engagement with an API-first communications startup in India. The product offered WhatsApp and SMS APIs through a 14-day trial. Some accounts evaluated independently; larger or more complex accounts involved sales during the same evaluation. The brief was more qualified pipeline without much more spend. The bigger constraint wasn’t volume alone. It was who was arriving, and what happened once they did.

Trial starts + demo requests (total recorded inbound acquisition volume)

down ~15%

down ~15%

Sales acceptance (rounded share of recorded trial/demo acquisition records later marked accepted by sales)

~1 in 5 → ~3 in 10

~1 in 5 → ~3 in 10

First-week activation among trial accounts

(completed and verified a representative transactional workflow)

~1 in 5 → ~1 in 3

~1 in 5 → ~1 in 3

Marketing-sourced qualified pipeline

(agreed use case, sufficient fit, a credible buying window, a next commercial step)

~100 → ~150

~100 → ~150

Paid cost per sales-accepted lead

~100 → ~80

~100 → ~80

Anonymized under NDA. The creative is reconstructed from engagement notes, not the client’s original work.

The brief was: grow qualified pipeline without spending a lot more to get it.

This was new to me. I hadn’t solved this kind of problem before. In my last job I’d been handed “grow the audience,” and I’d gotten reasonably good at it. So my first instinct was the obvious one, and it happened to be the thing I already knew how to do: find more of the right people, get more of them to raise a hand.

Then I looked at the funnel, and that instinct stopped making sense.

The company wasn’t short on interest. Traffic, campaigns, content, demo requests, a 14-day trial. The top was busy. That made me less sure that simply adding more activity would solve it. Before buying more traffic, I wanted to understand what was happening to the demand already coming in.

One distinction matters for the rest of the case. A trial start or demo request was an acquisition event. Activation applied only to trial accounts: completing and verifying a representative transactional workflow within the first week. Sales acceptance and qualified pipeline were commercial stages. The product and commercial paths could overlap — a trial account could later involve sales, and a sales-assisted account could still activate.

INBOUND ACQUISITION

Trial start

Demo request

PRODUCT EVALUATION

Trial account

↓

Representative transactional workflow

↓

Delivery verified in first week

↓

ACTIVATION

COMMERCIAL EVALUATION

When sales is involved

↓

Sales accepted

↓

Technical / commercial evaluation

↓

QUALIFIED PIPELINE

Paths can overlap

Activation measured product progress. Sales acceptance and qualified pipeline measured commercial progression.

The aggregate sales-acceptance rate is a rounded funnel measure reconstructed across trial/demo and CRM records. Account identity was reconciled after the fact rather than joined cleanly from the start.

The first-week window was deliberate. With a 14-day trial, reaching first value only near the end left little time to complete the rest of the evaluation before access stopped.

What was actually leaking

The thin pipeline could have come from a few different places.

Maybe the company wasn’t reaching enough relevant buyers. Maybe it was reaching plenty, but too many were the wrong kind of demand. Maybe suitable buyers were arriving and then struggling to understand the value or finish an evaluation. Or maybe the leak was later, in qualification and sales follow-up.

I kept those separate on purpose, because each one would have sent me somewhere completely different. One of them was to buy more traffic, which was also the direction I had the most experience with. So I wanted to see whether the evidence actually pointed there before going that way.

What the evidence said

I pulled from a few places: CRM and funnel data, trial behaviour, sales calls and notes, interviews with customers and prospects, stalled and lost evaluations, competitor pages, public reviews, and search behaviour.

I wasn’t trying to make every source answer every question. The behavioural and commercial data told me where accounts behaved differently. Interviews and sales conversations helped explain why. External research mostly helped me understand what buyers ran into before they ever reached us.

The clearest thing showed up only when I stopped averaging every lead into one funnel and looked at segment and use case together.

The first-week window was deliberate. With a 14-day trial, reaching first value only near the end left little time to complete the rest of the evaluation before access stopped.

What was actually leaking

The thin pipeline could have come from a few different places.

Maybe the company wasn’t reaching enough relevant buyers. Maybe it was reaching plenty, but too many were the wrong kind of demand. Maybe suitable buyers were arriving and then struggling to understand the value or finish an evaluation. Or maybe the leak was later, in qualification and sales follow-up.

I kept those separate on purpose, because each one would have sent me somewhere completely different. One of them was to buy more traffic, which was also the direction I had the most experience with. So I wanted to see whether the evidence actually pointed there before going that way.

What the evidence said

I pulled from a few places: CRM and funnel data, trial behaviour, sales calls and notes, interviews with customers and prospects, stalled and lost evaluations, competitor pages, public reviews, and search behaviour.

I wasn’t trying to make every source answer every question. The behavioural and commercial data told me where accounts behaved differently. Interviews and sales conversations helped explain why. External research mostly helped me understand what buyers ran into before they ever reached us.

The clearest thing showed up only when I stopped averaging every lead into one funnel and looked at segment and use case together.

The first-week window was deliberate. With a 14-day trial, reaching first value only near the end left little time to complete the rest of the evaluation before access stopped.

What was actually leaking

The thin pipeline could have come from a few different places.

Maybe the company wasn’t reaching enough relevant buyers. Maybe it was reaching plenty, but too many were the wrong kind of demand. Maybe suitable buyers were arriving and then struggling to understand the value or finish an evaluation. Or maybe the leak was later, in qualification and sales follow-up.

I kept those separate on purpose, because each one would have sent me somewhere completely different. One of them was to buy more traffic, which was also the direction I had the most experience with. So I wanted to see whether the evidence actually pointed there before going that way.

What the evidence said

I pulled from a few places: CRM and funnel data, trial behaviour, sales calls and notes, interviews with customers and prospects, stalled and lost evaluations, competitor pages, public reviews, and search behaviour.

I wasn’t trying to make every source answer every question. The behavioural and commercial data told me where accounts behaved differently. Interviews and sales conversations helped explain why. External research mostly helped me understand what buyers ran into before they ever reached us.

The clearest thing showed up only when I stopped averaging every lead into one funnel and looked at segment and use case together.

WHERE DEMAND CAME FROM

VS. WHERE PIPELINE CAME FROM

Share of demand (people raising a hand)

~25%

~27%

~48%

Share of qualified pipeline

~60%

~26%

~14%

Fintech & marketplace, transactional notifications

SaaS & commerce, mixed use cases

Very small, resellers, promotional, unclear

Rounded, anonymized figures. The priority segment was a quarter of demand and most of the pipeline; the long tail was the reverse.

A group that accounted for roughly a quarter of recorded trial starts and demo requests was producing around three-fifths of qualified pipeline. These were mostly mid-market fintech and marketplace accounts evaluating recurring, time-sensitive workflows: authentication, payment, payout, order and account-status notifications. Their trials also activated more often than the rest.

At the other end was a long tail. Very small businesses, resellers, promotional senders, vague use cases. Just under half of recorded trial starts and demo requests sat here, and it returned a low-teens share of pipeline, with something like one in ten trials activating. Lots of activity, very little of it commercial.

That was the first point where the brief changed shape for me. If we acquired more of the existing mix, we could hit a lead target and make the pipeline problem worse.

Even the best-fit segment mostly didn’t activate

This is the part I nearly flattened into a simpler story.

The priority group had the strongest activation of the lot. It was still only close to two in five trials. So “we found the ICP” wasn’t the whole answer, and I didn’t want to write it up as though it were.

So I split the comparison again, and kept three questions apart instead of collapsing them.

Why this group beat the rest was the easiest to see. For a team sending a delayed authentication, payment, payout or order notification, a failure isn’t abstract. It creates retries, support tickets, abandoned flows, manual investigation. They weren’t shopping for “customer engagement.” They had a workflow they needed to make work.

Why suitable accounts still didn’t activate took longer. Among otherwise promising trials, there was a recurring gap between starting setup and proving a realistic workflow. A useful evaluation meant more than calling an endpoint successfully. The team wanted to send a representative message, see its delivery state, understand what failure looked like, inspect enough to investigate it, and judge how it would fit into production. Some accounts never assembled that proof fast enough inside 14 days.

A group that accounted for roughly a quarter of recorded trial starts and demo requests was producing around three-fifths of qualified pipeline. These were mostly mid-market fintech and marketplace accounts evaluating recurring, time-sensitive workflows: authentication, payment, payout, order and account-status notifications. Their trials also activated more often than the rest.

At the other end was a long tail. Very small businesses, resellers, promotional senders, vague use cases. Just under half of recorded trial starts and demo requests sat here, and it returned a low-teens share of pipeline, with something like one in ten trials activating. Lots of activity, very little of it commercial.

That was the first point where the brief changed shape for me. If we acquired more of the existing mix, we could hit a lead target and make the pipeline problem worse.

Even the best-fit segment mostly didn’t activate

This is the part I nearly flattened into a simpler story.

The priority group had the strongest activation of the lot. It was still only close to two in five trials. So “we found the ICP” wasn’t the whole answer, and I didn’t want to write it up as though it were.

So I split the comparison again, and kept three questions apart instead of collapsing them.

Why this group beat the rest was the easiest to see. For a team sending a delayed authentication, payment, payout or order notification, a failure isn’t abstract. It creates retries, support tickets, abandoned flows, manual investigation. They weren’t shopping for “customer engagement.” They had a workflow they needed to make work.

Why suitable accounts still didn’t activate took longer. Among otherwise promising trials, there was a recurring gap between starting setup and proving a realistic workflow. A useful evaluation meant more than calling an endpoint successfully. The team wanted to send a representative message, see its delivery state, understand what failure looked like, inspect enough to investigate it, and judge how it would fit into production. Some accounts never assembled that proof fast enough inside 14 days.

A group that accounted for roughly a quarter of recorded trial starts and demo requests was producing around three-fifths of qualified pipeline. These were mostly mid-market fintech and marketplace accounts evaluating recurring, time-sensitive workflows: authentication, payment, payout, order and account-status notifications. Their trials also activated more often than the rest.

At the other end was a long tail. Very small businesses, resellers, promotional senders, vague use cases. Just under half of recorded trial starts and demo requests sat here, and it returned a low-teens share of pipeline, with something like one in ten trials activating. Lots of activity, very little of it commercial.

That was the first point where the brief changed shape for me. If we acquired more of the existing mix, we could hit a lead target and make the pipeline problem worse.

Even the best-fit segment mostly didn’t activate

This is the part I nearly flattened into a simpler story.

The priority group had the strongest activation of the lot. It was still only close to two in five trials. So “we found the ICP” wasn’t the whole answer, and I didn’t want to write it up as though it were.

So I split the comparison again, and kept three questions apart instead of collapsing them.

Why this group beat the rest was the easiest to see. For a team sending a delayed authentication, payment, payout or order notification, a failure isn’t abstract. It creates retries, support tickets, abandoned flows, manual investigation. They weren’t shopping for “customer engagement.” They had a workflow they needed to make work.

Why suitable accounts still didn’t activate took longer. Among otherwise promising trials, there was a recurring gap between starting setup and proving a realistic workflow. A useful evaluation meant more than calling an endpoint successfully. The team wanted to send a representative message, see its delivery state, understand what failure looked like, inspect enough to investigate it, and judge how it would fit into production. Some accounts never assembled that proof fast enough inside 14 days.

THE GAP THAT CAUGHT ACCOUNTS

The product could do all of this. Proving it took integration work, and the trial was running down.

MINUTES IN

202 Accepted

A test request went through.

“Does the API work?” Yes.

integration work

on the buyer’s side

14-day trial, running down

WHAT IT TOOK TO BELIEVE IT

Wire up delivery webhooks.

Trigger and inspect a failure.

Decide if it fits production.

Trust and commercial fit.

A successful send wasn’t a finished evaluation. Some accounts ran out of trial before finishing the work.

And activation still wasn’t the buying decision. Some accounts finished the technical test and stalled anyway. A successful request answered whether the API worked. It didn’t answer how hard this would be to implement properly, what happened when delivery failed, whether they could trust the vendor with a critical workflow, or what it would cost at production volume. That stopped me from treating one successful message as proof the customer was done evaluating.

The bet: fewer, better-fit accounts

For the next growth cycle, we concentrated on a specific buying situation:

  • Accounts: mid-market fintechs and marketplaces in India with in-house engineering.

  • Jobs: recurring authentication, payment, payout, order and account-status notifications.

  • Why they cared: a failed or delayed notification had a visible operational or customer consequence.

  • Why now: a near-term launch, rising volume, a reliability problem, channel expansion, or a vendor-consolidation decision.

The vertical label on its own wasn’t enough. A marketplace with high-volume order and payment messaging could be a better fit than a fintech using the product occasionally for promotions.

This was a growth-cycle priority, not a decision that the company should never sell to anyone else. New evidence could change it. For the test, we deliberately reduced effort aimed at generic engagement demand, promotional bulk messaging, rate-shopping reseller traffic, and very small accounts with no technical owner.

The message changed because the market choice changed

The original message was fine as category copy:

And activation still wasn’t the buying decision. Some accounts finished the technical test and stalled anyway. A successful request answered whether the API worked. It didn’t answer how hard this would be to implement properly, what happened when delivery failed, whether they could trust the vendor with a critical workflow, or what it would cost at production volume. That stopped me from treating one successful message as proof the customer was done evaluating.

The bet: fewer, better-fit accounts

For the next growth cycle, we concentrated on a specific buying situation:

  • Accounts: mid-market fintechs and marketplaces in India with in-house engineering.

  • Jobs: recurring authentication, payment, payout, order and account-status notifications.

  • Why they cared: a failed or delayed notification had a visible operational or customer consequence.

  • Why now: a near-term launch, rising volume, a reliability problem, channel expansion, or a vendor-consolidation decision.

The vertical label on its own wasn’t enough. A marketplace with high-volume order and payment messaging could be a better fit than a fintech using the product occasionally for promotions.

This was a growth-cycle priority, not a decision that the company should never sell to anyone else. New evidence could change it. For the test, we deliberately reduced effort aimed at generic engagement demand, promotional bulk messaging, rate-shopping reseller traffic, and very small accounts with no technical owner.

The message changed because the market choice changed

The original message was fine as category copy:

And activation still wasn’t the buying decision. Some accounts finished the technical test and stalled anyway. A successful request answered whether the API worked. It didn’t answer how hard this would be to implement properly, what happened when delivery failed, whether they could trust the vendor with a critical workflow, or what it would cost at production volume. That stopped me from treating one successful message as proof the customer was done evaluating.

The bet: fewer, better-fit accounts

For the next growth cycle, we concentrated on a specific buying situation:

  • Accounts: mid-market fintechs and marketplaces in India with in-house engineering.

  • Jobs: recurring authentication, payment, payout, order and account-status notifications.

  • Why they cared: a failed or delayed notification had a visible operational or customer consequence.

  • Why now: a near-term launch, rising volume, a reliability problem, channel expansion, or a vendor-consolidation decision.

The vertical label on its own wasn’t enough. A marketplace with high-volume order and payment messaging could be a better fit than a fintech using the product occasionally for promotions.

This was a growth-cycle priority, not a decision that the company should never sell to anyone else. New evidence could change it. For the test, we deliberately reduced effort aimed at generic engagement demand, promotional bulk messaging, rate-shopping reseller traffic, and very small accounts with no technical owner.

The message changed because the market choice changed

The original message was fine as category copy:

Build reliable customer conversations across WhatsApp and SMS.

Flexible APIs, delivery tools and one platform for customer communication as you scale.

Build reliable customer conversations across WhatsApp and SMS.

Flexible APIs, delivery tools and one platform for customer communication as you scale.

The problem wasn’t that it was badly written. It was that almost anyone shopping for communications software could see themselves in it.

The evidence had given us something more specific to say. A priority account wasn’t really asking for “better engagement.” It was trying to move a critical notification workflow into production without losing visibility when delivery went wrong.

That became the spine: critical workflow, production readiness, visible control when delivery fails. What it looked like depended on who was reading.

The person leading the evaluation varied by account. Sometimes engineering discovered and tested the product; in other accounts, product or operations led the evaluation while engineering proved the integration.

For an engineering evaluator, the point was that a successful send isn’t a successful delivery:

For a product or operations owner, the point was the gap between the thing that happened and the message about it:

For a product or operations owner, the point was the gap between the thing that happened and the message about it:

Same position, different entry points. The mechanisms underneath stayed concrete: API requests, delivery-state webhooks, message-level logs, failure investigation. That part mattered, because reliability on its own is category language. Plenty of credible providers can claim it. These showed what the buyer could actually see or do.

I’d landed on the transactional-reliability direction, but I wasn’t sure it was right yet, so I screened it against competing directions before we put weight behind it. Broad platform and channel coverage got steady response but kept mixing account quality and use cases. Cost and consolidation got the strongest attention and the cheapest forms, and also pulled in more rate-shopping and weak-fit interest.

Same position, different entry points. The mechanisms underneath stayed concrete: API requests, delivery-state webhooks, message-level logs, failure investigation. That part mattered, because reliability on its own is category language. Plenty of credible providers can claim it. These showed what the buyer could actually see or do.

I’d landed on the transactional-reliability direction, but I wasn’t sure it was right yet, so I screened it against competing directions before we put weight behind it. Broad platform and channel coverage got steady response but kept mixing account quality and use cases. Cost and consolidation got the strongest attention and the cheapest forms, and also pulled in more rate-shopping and weak-fit interest.

Transactional reliability and control got less indiscriminate response, but a better share of the priority accounts, sales-accepted leads and activated trials we actually cared about.

The first screen was a controlled paid-message comparison: audience, placement, budget rules and destination stayed as consistent as the platform allowed. Click-through and cost per form told us about attention and acquisition efficiency; sales acceptance and activation were the downstream quality checks.

Because copy and visual changed together in those screens, I treated them as directional rather than as proof that one sentence caused the difference.

For the stronger test, eligible visitors were randomly assigned to the existing priority page or the revised message-and-proof experience. The primary sales-path measure was visitor-to-sales-accepted conversion within the agreed attribution window; form completion stayed diagnostic. For the self-serve path, trial starts were followed into the first-week activation event. Raw form completion moved less than sales acceptance did, and that was fine. The point wasn’t to maximise the odds someone filled in a form. It was to improve the odds the person behind it was worth progressing.

Acquisition couldn’t make a promise the trial then dropped

Changing targeting and messaging wouldn’t have been enough if the right buyer entered the product and still had to reverse-engineer what “value” meant.

So I mapped the trial journey with the product and engineering counterpart, and reorganised it around the representative job: choose use case, create credentials, send a realistic request, verify delivery, inspect a failure state, identify who’d own it in production.

A user who’d never created credentials didn’t need the same lifecycle message as someone who’d sent a request and stalled at authentication. Someone who’d sent a message but never checked delivery state needed something else again. The intent wasn’t to manufacture more onboarding engagement. It was to shrink the distance between what we promised during acquisition and what the user had to prove before the evaluation meant anything.

High-fit or technically complex accounts could still get human help. But the success criterion stayed a verified product workflow, not the fact that a meeting happened.

“Not qualified” needed to mean something useful

We also tightened the feedback loop.

“Not qualified” isn’t useful when it can mean too small, reseller, promotional use case, no technical owner, no timeline, unclear use case, wrong geography, or no real volume. So the team agreed on a smaller set of mutually exclusive fit, intent and rejection definitions. Then we reviewed progression by segment, use case, source and message, instead of staring at one aggregate lead number.

The weekly question got deliberately simple: continue, revise, or stop? Higher activation was encouraging, but it stayed an early product signal while the paid and retained cohorts matured.

What changed, and by how much

We compared two similar, roughly two-month periods, before and after the main rollout. Media spend stayed broadly flat, and the stage and activation definitions were held stable.

  • Trial starts + demo requests: down roughly 15%.

  • Sales acceptance — rounded share of recorded trial/demo acquisition records later marked accepted by sales: from about 1 in 5 to about 3 in 10.

  • First-week activation among trial accounts: from about 1 in 5 to about 1 in 3.

  • Marketing-sourced qualified pipeline: from index 100 to about 150.

  • Paid cost per sales-accepted lead: from index 100 to about 80.

Trial starts and demo requests fell, and a report that stopped at that line would read it as a step back. But sales acceptance was higher on the reconstructed funnel measure, a larger share of trial accounts activated, and the paid cost of each accepted lead fell.

What I can claim, and what I can’t

Not every number above has the same causal strength.

The randomized page test is the cleanest evidence I have. The revised message-and-proof experience showed higher sales-accepted conversion for the tested traffic.

Activation is weaker. Onboarding guidance, lifecycle messaging and some product changes all overlapped during the before-and-after cohorts, so I can’t hand the movement to one of them.

Qualified pipeline sits further downstream again. Targeting, message, qualification, product help, sales execution and deal timing all fed into it. So I wouldn’t assign the roughly 50% increase to the programme alone. I can’t separate its share cleanly from the rest.

There’s also a plainer caveat. Two adjacent periods can differ for reasons that have nothing to do with me. And several paid and sales-assisted cohorts needed more time than the engagement allowed before I’d make serious retention or expansion claims. Those stayed as later validation.

What I owned

I owned the diagnosis and recommendation: framing the problem and competing hypotheses, analysing the funnel, CRM and segments, designing and synthesising the research, and turning that into the ICP recommendation, positioning and message architecture.

I also owned the customer-facing copy, experiment briefs and measurement definitions, the trial and lifecycle recommendations, and the weekly readouts.

The work itself was cross-functional. I worked with sales on fit and qualification, with product and engineering on the technical proof and trial journey, with design and web support on execution, and with the founder on prioritisation and the final calls.

What I’d change now

I’d define the account-level join before the first experiment, so anonymous visits, trial accounts and CRM records could be followed through the funnel instead of reconciled afterward.

I’d validate the activation definition against a larger paid and retained cohort. Completing a verified transactional workflow was meaningful behaviour, but I’d want to know whether it actually predicted durable usage rather than just looking like an important milestone.

I’d recruit differently. My sample leaned toward customers and active prospects, the people who could explain why an evaluation worked. I’d want more people who never activated or chose another vendor, because they could tell me more about where an evaluation quietly died.

And I’d preserve a longer holdout around the trial changes. The page experiment isolated one decision reasonably well. The activation work was harder to separate because several things changed together. I learned from it, but I’d be more cautious about what I attributed to any one change.

What stayed with me

I came in with more than two years around CPaaS products, so I already knew the category and most of its buying language. That helped me ask questions sooner. It didn’t answer them.

The useful shift came when we stopped treating every signal of interest as progress, and started separating a few things we’d been blurring together: attention, intent, product value, commercial fit, pipeline, and durable customer value.

By the end, the number that first looked worst was the 15% drop in trial starts and demo requests. Once those things stopped being one funnel, I stopped reading that decline the way I would have at the start.

How I got to this (methods)

A short note for anyone who wants the how, not just the what.

  • Segment concentration, ranked by value. I ranked segment-and-use-case combinations by the qualified pipeline they produced, instead of treating every lead as equal. That surfaced the pattern where roughly a quarter of recorded trial starts and demo requests accounted for around three-fifths of qualified pipeline.

  • Cohort analysis. I compared accounts by when they entered, so I wasn’t judging a two-week-old trial against a two-month-old one.

  • Competing hypotheses. Before choosing a direction, I held the four explanations for the weak pipeline side by side and let the evidence decide which mattered most.

  • Randomized A/B test. For the message-and-proof page, eligible visitors were randomly assigned to the old or new experience. That’s the result with the cleanest attribution: the new experience showed higher sales-accepted conversion.

  • Qualitative coding of calls and interviews. To explain why segments behaved differently, not just that they did.

© Sanjana A P | all rights reserved | 2026

© Sanjana A P | all rights reserved | 2026