Maximum collection control
Keep your research collectors and evidence logic.
Best for teams with mature collectors and their own research-data operation.
Explore proxy infrastructureRésultats de recherche sur le marché et la concurrence
Turn public-web sources into traceable market and competitor data through structured APIs, scheduled feeds, or a managed research program.
Start with the workload
Choose the question. We define the public evidence, coverage, and limits needed to support it.
Evidence-base assembly
Collect observable organizations, offers, locations, and activity; your team applies the assumptions behind any estimate.
Research method
Set the entities, discovery rules, source coverage, markets, and blind spots that make the work repeatable.
Set approved entities and candidate-review rules.
Set markets, segments, source types, and thresholds.
Preserve locale, query, postcode, currency, and result settings.
Record unsupported sources, fields, and method limits.
Evidence model & data contract
Representative records confirm coverage, fields, provenance, and quality rules before production.
We test source eligibility, fields, market context, cadence, and feasibility.
Source eligibility, field availability, and collection conditions are confirmed during feasibility review. Restricted sources are not included.
Source ID, canonical ID when included, name, aliases, domain, location, and review state.
Source text or value, page type, record type, and separately labeled normalized fields.
Source URL, class, market, locale, query, currency, and other result-shaping context.
Captured at, published at, first seen, last seen, state, and previous value.
Observed, not exposed, not observed, failed, outside scope, removed, unresolved, or review required.
// Illustrative managed record with entity resolution included { "entity": { "source_id": "company:source:1842", "canonical_id": "company:wsa:00917", "match_state": "reviewed" }, "source_observation": { "page_type": "pricing", "field": "plan_name", "raw_value": "Enterprise", "normalized_value": "enterprise" }, "source_url": "https://example.com/pricing", "market": "GB", "captured_at": "2026-07-14T06:15:00Z", "derived_event": { "type": "offer_added_candidate", "comparison_window": "2026-07-07/2026-07-14", "rule": "not_observed → observed", "review_state": "pending" } }
# Illustrative research specification question: "Track competitor expansion signals" universe: seed_entities: 42 discovery: "approved directories → review queue" markets: ["GB", "DE", "FR"] sources: - "official company pages" - "career pages" - "business location listings" cadence: "weekly collection window" history: "snapshots + first_seen + last_seen when scoped" evidence_retention: "source URL; retained capture only when contracted" delivery: "JSONL to agreed destination" known_blind_spots: - "private transactions" - "offline activity" - "sources outside approved universe"
// Illustrative acceptance rules coverage: planned_sources_reported: true collection_errors_separate: true provenance: source_url_required: true captured_at_required: true context_fields_required: "by record type" entities: unresolved_state_allowed: true force_uncertain_match: false history: compare_only_like_context: true missing_means_absent: false delivery: collection_window: "source-tested and contracted" required_field_fill: "set from representative sample" late_or_failed_run: "report + agreed recovery policy" schema_change_notice: "agreed notice window" maintenance: assigned_to: "per service model" scope_change_requires_review: true
Quality & longitudinal evidence
Report planned, collected, unavailable, and failed records.
Validate required fields, types, accepted values, and source constraints.
Preserve exact, probable, unresolved, and reviewed states when included.
Keep not exposed, not observed, failed, removed, and outside scope distinct.
Retain source URL, context, timestamps, schema, and comparable history.
WSA monitors and maintains every supported endpoint and connector it operates.
Operating model
Maximum collection control
Best for teams with mature collectors and their own research-data operation.
Explore proxy infrastructureScheduled & managed engagement
Confirm the question, sources, markets, entities, and fields.
Review coverage, context, and representative records.
Set the schema, cadence, quality rules, and responsibilities.
WSA launches and maintains the contracted workflow.
Your evaluation package includes representative records, coverage notes, a draft schema, and the proposed delivery plan.
Evaluation questions
Coverage, evidence, history, ownership, and commercial terms—answered clearly.
WebScrapingAPI provides public-web data infrastructure and delivery, not research or BI software. Your team owns interpretation and decisions.
For a managed program, we can turn your question into sources, entities, inclusion rules, fields, cadence, and blind spots. The result defines observable coverage, not the entire market.
Yes, when discovery is included. Candidates from approved public sources enter a review state before joining the tracked set.
Entity resolution remains with your team for proxies, the Scraper API, structured APIs, and standard feeds. It can be added to a feed or managed program, where WSA uses approved rules and review states.
We keep not observed, not exposed, failed, outside scope, and verified absence distinct. Conflicting records retain their own source and capture time.
Structured APIs return current observations. Feeds and managed programs can add snapshots, first-seen and last-seen fields, and change records; history begins when collection starts unless a backfill is confirmed.
WSA maintains retrieval for the Scraper API, access, extraction, output monitoring, and source changes for structured APIs, and every connector it operates for feeds and managed programs. Proxy customers maintain their collectors; Scraper API customers maintain parsing and downstream logic.
Eligible public pages can be evaluated for a feed or managed program; restricted sources are not included. WSA delivers the selected public text and metadata, while your team derives sentiment, intent, and market conclusions.
Self-serve products follow published plans. Custom pricing reflects sources, entities, markets, discovery, cadence, fields, resolution, history, validation, format, and delivery.
Review our Politique de confidentialité, GDPR information, and Accord de service. Custom agreements define source eligibility, data handling, retention, access, responsibilities, change handling, and source constraints.
Validate the research brief
Share your question and representative sources. We will validate feasibility and propose the data model, cadence, and delivery plan.