Blog

Executing actions during web scraping: a practical guide

September 28, 2026

dynamic web scraping, Web Scraper selectors, data quality, Browser actions, Pagination

A page can load successfully and still show the wrong data. A product price may depend on the selected size, a directory may need a location, and a job board may reveal more listings only after a click. Executing actions during web scraping means reproducing the steps that create the specific page state your dataset requires.

The useful sequence is action → expected state → readiness → extraction → record check. A click that succeeds without producing the intended records is not a successful scraping step.

Continue Reading

Web scraping API vs Web Scraper Cloud: How to choose

September 28, 2026

data extraction, Web Scraper Cloud, scraping costs, data quality, Web scraping API

A scraping API, such as ScrapingBee's HTML API, fetches a page when your application sends a request. Web Scraper Cloud runs a tested sitemap that defines how to navigate a site and assemble records across pages. Both can use browser rendering and proxies. The practical difference is how much of the collection workflow your team builds and maintains.

For a few known URLs inside an existing application, an API can be simpler. For a recurring catalogue, directory or marketplace dataset, a reusable sitemap, hosted schedule and checks on the finished data can remove substantial work.

Continue Reading

Web Scraping Pricing: Capacity, requests and credits

September 25, 2026

web scraping pricing, Web Scraper Cloud, scraping APIs, scraping capacity

Web scraping platforms sell two quite different things: capacity to run scrapers or a quantity of billable work. Web Scraper Scale charges for concurrent scraper capacity with unlimited URL credits. Many scraping APIs instead deduct requests, credits or delivered records from a monthly allowance. A low price per request says little until you know which configuration that request needs.

This guide separates those models first, then shows how Web Scraper, Decodo, ScrapingBee and other platforms apply them. It also explains why the amount of data you receive and the time needed to collect it matter as much as the published rate.

Continue Reading

Instant Data Scraper can collect browsing, AI and shopping data if you opt in

September 23, 2026

data security, extension privacy, Instant Data Scraper, code review, web scraping, Chrome extensions

Instant Data Scraper has changed hands: Web Robots, which created and formerly owned the extension, says it no longer owns or supports it. The Chrome Web Store now lists Flavr Technology LP as its publisher, with one million users. It still exports tables, but version 1.7.1 also contains code that watches page visits and runs rules targeting AI chats, Amazon cart contents at checkout, search results and ads.

For an opted-in user, selected content can be queued and attached to a later browsing event sent to the publisher. Flavr’s policy describes collection and affiliate sharing for extension users generally, including those outside the US and Europe. Sections of that same policy for California and other US states explicitly discuss selling or sharing these data categories. We did not verify a sale of an individual user’s data.

This collection is separate from the table you deliberately export. The reviewed code does not show that exported table being uploaded, and its main page-content upload requires opt-in. Analytics and configuration requests can still run if you decline. Start with what this means for users; the code, network paths and reproduction scripts follow below.

Continue Reading

Marketplace vs custom sitemap: which should you use?

September 22, 2026

Web Scraper Cloud, data quality, web scraping, custom sitemaps, Sitemap Marketplace

Choose a Marketplace sitemap when a ready-made scraper covers your target site, accepts the pages you need to start from and returns the fields your dataset requires. Build a custom sitemap when the source, required fields or navigation differ. Marketplace sitemaps are not editable: changing the input URLs does not change their selectors, page paths or extracted fields.

Here, a sitemap means Web Scraper's instructions for navigating pages and extracting records. It is different from an XML sitemap a website publishes for search engines.

Continue Reading

How often should you schedule a web scraper

September 22, 2026

data freshness, Web Scraper Cloud, data quality, Scheduled web scraping, Web scraping strategy

Schedule a web scraper often enough to deliver useful, valid data before someone needs to act on it. A promotion ending this afternoon may justify several checks during the day; a directory used for monthly planning may need far fewer. Neither hourly nor daily scraping is the right default for every source.

Choose a starting cadence from the decision deadline and the changes you need to catch. Then test whether the scraper, the website and your data checks can sustain it.

Continue Reading

How to track promotions across online stores

September 21, 2026

e-commerce data, promotion monitoring, competitor monitoring, web scraping, data automation

To track promotions across online stores, check the pages where stores announce sales and the product pages where an offer applies. Save the wording, prices, eligible items, conditions, market and time observed. Repeat the checks, validate each collection, then compare the results to flag newly observed, changed and apparently ended offers.

A price drop is only one kind of promotion. Coupons, multi-buy deals, member prices and free delivery each have different rules. A useful monitor tells your team what the store offered and what remains unverified, rather than reporting every change as a percentage discount.

Continue Reading

Resume interrupted scraping jobs without duplicates

September 20, 2026

idempotency, retries, data quality, data pipelines, web scraping, Scraping architecture

To resume interrupted scraping jobs safely, do not continue from a worker's last loop index. Reconstruct unfinished work from durable task state, assume that a task may run more than once and make repeated writes harmless with stable keys and atomic checkpoints.

This guide focuses on recovery mechanics: task states, leases, transaction boundaries, retries and targeted reruns. For broader decisions about capacity, concurrency and URL frontiers, see scaling web scraping from thousands to millions of pages.

Continue Reading

Cloud vs local web scraping: How to choose

September 18, 2026

web scraping infrastructure, automation, Web Scraper Cloud, local web scraping, Cloud web scraping

A scraper’s interface does not tell you where its jobs run. A browser extension can create workflows that later execute in the cloud, while a Python scraper can run on a laptop, an office server, or a cloud virtual machine.

The practical decision is therefore not simply “local or cloud.” It is who operates the execution environment, keeps it available, manages browser capacity, handles failures, and delivers the resulting data.

Local execution is usually the simplest starting point. Cloud execution becomes more useful when the dataset must refresh without supervision, meet a delivery deadline, or serve more than one person. Self-managed and managed cloud systems achieve that in different ways.

Continue Reading

Batch vs continuous web data collection

September 18, 2026

data freshness, Web Scraper Cloud, data pipelines, Scraping architecture

Batch and continuous web data collection solve different freshness and completeness problems. Batch collection processes a defined scope and publishes a validated snapshot. Continuous collection keeps a persistent queue running and publishes smaller updates as individual targets become due.

Neither model is universally faster, cheaper, or more reliable. The right choice depends on how fresh the data must be, whether consumers require a complete point-in-time snapshot, and how much operational complexity the pipeline can support.

Continue Reading

When browser rendering becomes a scaling bottleneck

September 17, 2026

headless browsers, Web Scraper Cloud, Scaling, Web scraping performance, FullJS

Browser rendering becomes the scaling bottleneck when the browser stage cannot turn queued URLs into validated pages as quickly as the rest of the scraping pipeline can supply and process them. At that point, adding more work increases queueing, page latency, retries or failures without producing a proportional increase in accepted data.

The solution is not automatically more browser concurrency. First prove which stage is saturated, calculate capacity from accepted output and remove browser work that the dataset does not require.

Continue Reading

How to scrape conference exhibitor lists accurately

September 16, 2026

market research, data quality, event intelligence, lead generation

To scrape conference exhibitor lists reliably, first map how the public directory works, define an event-specific schema, build one record for each exhibitor listing, follow profile links for additional fields, and make sure every page or dynamically loaded result is covered. Only automate the sitemap after a limited test produces a complete, correctly structured export.

The difficult part is rarely extracting a company name. It is preserving the relationship between the exhibitor, its booth, categories, profile, and event edition without missing records or deleting legitimate duplicates.

Continue Reading

Extract website data to Excel or CSV automatically

September 16, 2026

Excel, Web scraping automation, data export, CSV

To extract website data to Excel or CSV automatically, define the columns you need, build a scraper that reaches every relevant record, validate the result, and then schedule the completed dataset for delivery. The file format is only the final step. Reliable automation also depends on navigation, consistent field structure, and quality checks.

For a simple public table, Excel Power Query may be enough. For recurring product, property, job, marketplace, or directory data spread across listing and detail pages, a visual scraper with remote execution is usually more practical.

Continue Reading

Build a price history dataset from scheduled scrapes

September 16, 2026

Time-series data, e-commerce data, data pipelines, Scheduled web scraping

A scheduled scraper gives you repeated price snapshots. It does not create a reliable price history by itself. To build one, you need to preserve when each price was observed, connect observations to a stable product or offer, record gaps honestly and prevent failed or duplicate imports from becoming market data.

This guide shows how to turn recurring scrapes of a known set of public product pages into an auditable time-series dataset. Web Scraper Cloud handles tested extraction, scheduling and delivery, while your downstream system owns historical storage, acceptance rules and analysis.

Continue Reading

How to automate Web Scraper Cloud with the API

September 15, 2026

Web scraping automation, Web Scraper Cloud, data quality, data pipelines, API

To automate Web Scraper Cloud with the API, treat every scrape as an asynchronous data-pipeline run: identify the sitemap, create a job, record its ID, wait for a terminal status, download the result, validate it, and only then publish it downstream. The API removes the need to start and monitor recurring jobs manually, while Web Scraper Cloud handles the browser execution and collection work.

This tutorial builds that workflow in Python. It also covers the parts that make an automation safe in production: rate limits, ambiguous POST outcomes, newline-delimited JSON, data-quality checks, quarantine, and atomic promotion.

Continue Reading