Blog/Buying signals

Build a Lead Scraper, or Buy the Panel

A lead scraper takes a weekend to build and never stops needing attention. The build cost is not what decides it — the maintenance is, and so is whether you need something we deliberately do not offer.

Nils SpölgenSeptember 4, 202616 min
Who this is for

Agencies and software companies deciding whether to build their own ecommerce store scraper or buy a maintained panel

TL;DR
  • →Writing a lead scraper is the cheap part. Keeping one accurate is a standing job: markup changes, anti-bot measures, and a re-scrape schedule that has to run whether or not anyone is watching it.
  • →The build-versus-buy question is not about the code. It is about who absorbs the maintenance, and how stale your data is allowed to get before it stops being worth sending against.
  • →Keaz Signals runs that panel: 4.1 million Shopify and WooCommerce stores across Europe and North America, re-scraped weekly, as of August 2026. Two platforms, two regions, deliberately.
  • →Building your own wins outright in four cases — a platform we do not cover, a market outside the EU and North America, a field we do not collect, or data you need in your own warehouse. We do not export, so that last one is a real loss for us.
  • →We are not publishing a scraping success rate, a breakage frequency, or a percentage of stores refreshed on time. We did not measure them, so the numbers do not exist here.

What a lead scraper is, and where the cost actually sits

A lead scraper is a program that visits business websites and pulls structured fields off them: the platform a shop runs on, the apps it has installed, its contact addresses, its social accounts. Writing one is a weekend. Keeping one accurate is a standing job, and that asymmetry is the whole of the build-versus-buy decision.

The disclosure belongs up here, because it should change how you read the rest. We build Keaz Signals, which sells ecommerce leads with live buying signals attached. We run the panel this post says you might buy instead of building. We are a competitor to your own engineering time, and we have an interest in how you weigh it.

So this is not an argument that you should not build one. Plenty of people should, and the section on where building wins is not a courtesy: it names four cases where we are the wrong purchase, and one of them is a straightforward loss for us. What this post argues is narrower. The build cost is the number everyone estimates. The maintenance cost is the number that actually decides it.

It also helps to be clear about which job a scraper is doing. It is the finding layer, building the pool and filtering it down. It is not contact verification, not sending, and not the decision about who to contact this week. We have written separately about the four jobs that hide inside a prospecting stack. A scraper answers one of them, and building it does not make the other three appear.

Why the first version works and the third month does not

The first version of a scraper works because you wrote it against the pages you were looking at that week. The third month is harder for reasons that have nothing to do with how well you wrote it. Shops change their markup, hosts add bot protection, and a schedule that nobody owns quietly stops running.

Four things generate the ongoing cost, and they are structural rather than a comment on anybody's engineering.

  1. Markup drift. A shop's storefront is a product that other people are actively redesigning. Every theme update, app install and platform release is a chance that the selector you relied on now points at nothing. Nothing warns you; the field just comes back empty, and an empty field looks the same as a shop that genuinely does not have the thing.
  2. Access. Rate limits, bot detection and challenge pages are not a one-off problem to be solved but a moving one, because the other side is also maintaining their side. Whatever you build to get past today is a thing you now own and keep.
  3. Scheduling. A pool is only as good as its refresh cadence, and cadence is infrastructure: queues, retries, partial failures, and someone noticing when last night's run covered a fifth of what it should have. This is where most home-built pools decay, because the decay is invisible from the inside. The rows are still there. They are just older than you think.
  4. Coverage discovery. Scraping a list of shops you already know about is a different problem from finding the shops you do not. A pool that only ever re-reads its own contents stops growing on the day you stop hand-feeding it.

None of that is an argument that the work is hard. It is an argument that the work never finishes, which is a different property, and the one that matters when you are pricing the decision. A build is a project with an end. A panel is a subscription with an end. A build that has to stay current is a subscription you pay in your own team's attention, and it is the only one of the three that does not appear on any invoice.

What building it yourself genuinely gets you

Building your own gets you three things no vendor can sell you, and an honest version of this argument starts there rather than with the maintenance bill. You get exactly the fields you asked for, you get them in your own systems, and you get to point the thing at whatever corner of the market you like.

  • Fit. Every panel is somebody else's opinion about which fields matter. If your qualification depends on something idiosyncratic, whether a shop offers subscriptions, which payment methods appear at checkout, how many reviews sit on a product page, a scraper you wrote collects that and a panel you rented probably does not.
  • Ownership. The data lands in your warehouse, joins to your other tables, and stays after you stop paying anyone. If your business is going to be built on this data over years, that is a serious argument and it gets stronger the longer the horizon.
  • Reach. You can aim it anywhere. Platforms nobody sells a panel for, marketplaces, regional store networks, niches too small for a vendor to bother with. Coverage is a commercial decision for a vendor and a technical one for you, and yours can be stranger than theirs.

There is a fourth advantage that gets less credit. Building it teaches you what the data actually is. Teams that have scraped their own market once tend to be much better buyers afterwards, because they know which fields are cheap to collect and which are hard, and they stop being impressed by the cheap ones. If you are choosing between the two options and have never done it, the exercise has value beyond its output.

So the maintenance argument is a cost to be weighed, not a trump card. Plenty of teams look at that cost, decide it is affordable, and are right.

A scraper gives you states; the decision needs events

A scraper naturally produces states, and a cold email needs an event. A state is a durable fact: this shop runs Shopify, sells apparel, sits in Belgium, has a particular email tool installed. An event is a dated fact: it installed that tool eleven days ago. The second one tells you when to write. The first one does not.

This is the part of the build-versus-buy question that gets missed, because it is not a question about data quality. You can scrape states perfectly and still be no closer to knowing which forty of your nine thousand matching shops are worth a message this week.

Two consequences follow. The first is that a filter over a static pool is a filter anyone else can run. If four agencies sell the same service into the same segment and describe that segment the same way, all four are looking at the same shops on the same Monday. A state field is not scarce, so it cannot be the differentiator.

The second is that events only exist if you were watching before they happened. A single scrape gives you a snapshot, and a snapshot has no dates in it. To turn states into events you need the same shops read repeatedly on a schedule and the differences stored, which means the interesting output of a scraping operation is not the scraper at all. It is the history. That is a much larger thing to own than the code, and it cannot be backfilled: whatever you did not record last month is simply gone.

That is a structural difference between a snapshot and a series, not a claim that anyone's scraped data is dishonest. A snapshot is exactly the right tool for describing a market. It is being asked, by whoever built it, to answer a question about timing that its shape does not contain.

If that distinction is new, the longer version is in our write-up of what a sales intelligence platform knows and what it cannot, which makes the same argument about bought data rather than built data.

What we run instead, and what it is not

What we run is the maintained version of the thing described above, for one vertical. The pool is 4.1 million Shopify and WooCommerce shops across Europe and North America, re-scraped weekly, as of August 2026. Those numbers move, so treat the date as part of the number.

Two platforms and two regions is a small pool by the standards of a general lead database, and the narrowness is deliberate rather than a stage we are growing out of. We picked the vertical where the events are visible from outside. A shop is a public thing: its ads run somewhere you can look, its storefront changes visibly, its product pages appear and disappear, its social accounts carry dates. Reading a shop every week produces events. Reading a private company's org chart mostly produces a slightly newer version of the same state.

Because the weekly read is stored rather than overwritten, the record is a shop plus what changed about it recently: new Meta ads running, active ad count rising, a product launch, an email marketing tool installed, social growth, newsletter activity, storefront and site changes. The full list of what we currently detect is the buying signal catalogue, and what sits underneath it is described on the ecommerce leads database page. Read both before a trial, because they are also the honest place to find out that no signal we detect matches what you sell.

Contacts arrive with the shop, including founder addresses that are not publicly listed, so the enrichment step that usually follows a scrape is not a separate build. Segments are assembled with country, platform, follower and other conditions, with a live audience estimate as you narrow them. Sending runs through your own Instantly workspace, so the domains and mailboxes stay yours and replies sync back, and the sales agent drafts per-shop copy from the signal plus a knowledge base you write once. Access is capped per market, which is the mechanism that stops the segment you buy being sold to four of your competitors in the same week.

Several things are not there and are worth knowing before you weigh the trade. Registry and firmographic triggers and hiring signals are marked planned on our own site, which means they do not exist today. There is no CSV export. There are no LinkedIn signals and no LinkedIn sending, and no plans to add either. On data protection our position is deliberately narrow: we are GDPR-conscious by design, and responsibility for what you send stays with you. We will not claim more than that. That is a statement about where the responsibility sits, not a claim about anyone else's compliance.

Where building your own scraper is the better choice

Four cases make building your own the better decision, and in each one we are not a close second. If any of them describes you, write the scraper and do not spend another afternoon on comparison posts.

  1. You need a platform we do not cover. Ours is Shopify and WooCommerce. Magento, BigCommerce, Wix, Squarespace, PrestaShop, Shopware, headless builds, marketplace sellers with no storefront of their own: none of that is in our pool, and we are not going to present a thin slice of it as coverage.
  2. Your market is outside Europe and North America. If you sell into Latin America, Asia, Africa or Oceania, we do not have those shops. A scraper you point yourself has no such boundary, and this is exactly the kind of gap a vendor draws for commercial reasons and you do not have to.
  3. You need a field we do not collect. Our fields are the ones that feed the signals we detect. If your qualification turns on something else entirely, you either persuade a vendor to add it, which is slow and usually fails, or you collect it yourself, which is a build.
  4. You need the raw data in your own warehouse. We do not export. Leads move into campaigns and stay there. If your CRM or your data platform is the system of record, or you are building a product on top of this data, that rules us out and it should. This is a real loss for us in this comparison and there is no version of it where we win the point.

A fifth case is less about capability and more about arithmetic. If your total addressable market is a few thousand shops, the whole question is smaller than it looks. You can read a few thousand shops by hand over a couple of weeks, or with a scraper that never has to be robust, and neither a panel nor a serious build is warranted. Build-versus-buy only starts to matter at a scale where nobody can keep the pool current by attention alone.

There is also a limit that applies to both options equally. More rows and faster drafting do not create capacity to answer replies. The clearest sign that either approach is worth the money is paid sales capacity that is not fully booked. If your calendar is filled by referrals, you probably do not need one yet.

What we are not going to tell you about scraping

A post like this is supposed to contain a table of numbers proving the panel wins. There is not going to be one, and the reason belongs in the text rather than in a footnote at the bottom.

We are not going to tell you how often a home-built scraper breaks, because we did not measure that and neither has anyone else writing about it. The figure would have to come from a population of other people's scrapers that nobody has ever assembled. Any number in that shape is an invention with a decimal point on it.

We are also not going to publish a scraping success rate, a field-level accuracy rate, or a percentage of our own pool refreshed on schedule. We could design a test for the last one and we have not run it, so the number does not exist here. Stating it anyway would make this post feel more rigorous and be less true, and once a figure like that is in print it gets quoted back for years.

The same restraint rules out the comparison you might expect next. We are not putting a build cost next to a subscription price, because what each side meters is not the same thing. One is engineering hours you already pay for, spread over an unknown number of years; the other is a per-market cap on access. A table that lines those up as if they were comparable would be arithmetic dressed up as a service to the reader.

And nothing here is a promise about outcomes. We will not quote a reply rate or a deliverability result for a signal-led send, because a reply rate is a property of your offer, sent to your list, from your sending setup. A vendor controls one of those three. A figure produced under someone else's conditions is not evidence about yours.

What survives once the unmeasured claims are stripped out is the structural argument: a snapshot is not a series, and a series has to be maintained by somebody. You can check that yourself in an afternoon, which is more than you can do with a table of numbers. Our head-to-head write-ups against named alternatives, including the ones where we recommend the other product, live on the comparison pages rather than in blog posts, so they can be kept current in one place. If the option you are weighing is a directory of shops rather than a build, Keaz Signals compared with store directories is the relevant one.

How to decide this in an afternoon

You can settle this without trusting anyone's published figures, including ours. Five questions decide it, and all five are answerable from your own situation in an afternoon.

  1. Write down the platforms and countries you actually sell into. If any of them sits outside Shopify and WooCommerce in Europe and North America, the answer is already build, and the rest of the questions are academic.
  2. List the fields your qualification genuinely turns on, then check them against what a panel publishes. Not the fields that would be nice, the ones that decide whether a shop is worth contacting. If two of them are missing, you are building whatever you buy.
  3. Ask where the data has to live afterwards. If the answer is your warehouse or your CRM, stop here and build, because we do not export and that is not negotiable.
  4. Name the person who will own the refresh in six months, by name, and the hours a week. If nobody can be named, you are not choosing between a build and a panel. You are choosing between a panel and a pool that quietly goes stale.
  5. Decide whether you need states or events. If a filtered list is enough because your offer is relevant whenever someone reads it, a modest scraper does the job and the history does not matter. If your message only lands when something just happened, you need a series, and a series has to have been running before today.

If those answers point at buying, and you sell to Shopify or WooCommerce shops in Europe or North America, we are worth twenty minutes. There are 1,000 free leads on signup, which is enough to find out whether your segment exists in our pool and whether a signal fires in it often enough to matter, before any money changes hands. What it costs after that is on the pricing page, and the pool itself is described on the ecommerce leads database page. We would rather you test it against your own segment than take our account of the trade-off on trust.

Sources

  • Keaz Signals product pages, retrieved 4 September 2026: /ecommerce-leads-database, /signals, /sales-agent, /compare and /pricing. Source for the signal list, the segment builder, the per-market access cap, the Instantly sending model, the export limitation, the LinkedIn position and the triggers marked planned rather than shipped.
  • Our own pool figures: 4.1 million Shopify and WooCommerce shops under watch across Europe and North America, re-scraped weekly, as of August 2026. These numbers move; the date is part of the number.
  • 1,000 free leads on signup, as of August 2026.
  • Related posts on this blog: what sales prospecting tools actually do, what a sales intelligence platform knows and agency lead generation when your buyers run stores.
  • No scraping success rate, breakage frequency, field accuracy rate, refresh-completion percentage, competitor pool size, competitor freshness figure, price comparison or reply rate is cited anywhere in this article. That is deliberate, and the reason is in the section on what we are not going to tell you. We did not run those measurements, so we do not report them.

Questions we get

Is it cheaper to build a lead scraper or buy a lead database?

The build is almost always cheaper to start and the comparison is not really about that. A scraper you write is a one-off cost plus an open-ended maintenance cost: markup changes, bot protection, a refresh schedule someone has to own, and coverage discovery for shops you do not yet know about. A panel converts that into a subscription. Which is cheaper depends on how long your horizon is and whether you can name the person who will keep the scraper current in six months.

How hard is it to keep a lead scraper working?

The difficulty is not in any single fix; it is that the work never finishes. Storefronts are actively redesigned by other people, so selectors break without warning and a broken selector returns an empty field that looks identical to a shop which genuinely lacks the thing. Access controls move too, because the other side is maintaining their side. We have not measured how often other people's scrapers break and we are not going to publish a figure for it.

What can a scraper not tell you?

When to send. A scrape produces states, which are durable facts: this shop runs Shopify, sells apparel, has a particular email tool installed. A cold message needs an event, which is a dated fact: it installed that tool eleven days ago. Turning states into events requires the same shops read repeatedly on a schedule with the differences stored, so the valuable output of a scraping operation is the history rather than the code. History cannot be backfilled.

Can I export leads from Keaz Signals into my own warehouse?

No. We do not export; leads move into campaigns and stay there. If your CRM or data platform has to be the system of record, or you are building a product on top of the data, that rules us out and building your own is the better choice. It is the clearest case in this comparison where we lose outright.

Which ecommerce platforms and regions does Keaz Signals cover?

Shopify and WooCommerce shops in Europe and North America, 4.1 million of them re-scraped weekly, as of August 2026. Magento, BigCommerce, Wix, Squarespace, Shopware, headless builds and marketplace-only sellers are not in the pool, and neither are markets outside those two regions. If your buyers sit outside that boundary, a scraper you point yourself has no such limit.

Nils Spölgen
Co-founder · Keaz

Builds the signal pipeline behind Keaz Signals. Writes about what the store data actually supports, and what it does not.

Keep reading

Reading about signals is fine. Seeing yours is better.

Access opens per market, in order of signup. When your seat is ready you see which stores in your niche are moving and what we would send them.

Launch my agent