Aller au contenu

Ensembles de données et flux

Data Marketplace for ready-to-use web datasets.

Browse ready-to-use structured datasets by domain and source. Inspect the schema and representative sample records, then select the collection that fits your analysis.

  • Find the source Domain, entities, markets
  • Inspect the records Fields, types, missing states
  • Choose the timeline Snapshot and available history
  • Plan the handoff Format, partitions, destination

Dataset catalog

Browse datasets by source and record type.

Start with the entities your team needs, then open a dataset page to review fields, applications, refresh options, and a representative record.

01 · record

Start with the entity

Separate products from offers, companies from locations, and properties from listings before comparing fields.

Object and use case

02 · sample

Test fields with real context

Review identifiers, source context, timestamps, optional values, and missing states in representative rows.

Schema and joins

03 · next step

Choose the right delivery path

Select a prepared snapshot, or continue to Data Feeds when the collection needs a recurring schedule and destination.

Snapshot or recurring feed

Schema and fields

Understand every record before it enters your pipeline.

Use the data dictionary, source context, identifiers, timestamps, and missing states to verify that each field supports the joins, filters, and calculations your team will run.

Résumé de l'échantillonJSON illustratif
{
  "product_id": "wsa_20491",
  "title": "Trail shoe · blue · 42",
  "offer": {
    "price": 119.00,
    "currency": "EUR",
    "availability": "in_stock"
  },
  "source_url": "https://example.test/item",
  "captured_at": "2026-08-05T09:30:00Z",
  "schema_version": "commerce.v1"
}
A representative structure connecting product, offer, source, and collection time.
Révision du dictionnaire des donnéesCe qu'il faut vérifier
Entity identity
Stable identifiers, source identifiers, matching keys, and parent-child relationships.
Observation context
Source, market, requested context, capture time, and schema version.
Value semantics
Data type, unit, currency, normalized value, and source value where relevant.
Missing states
The difference between absent, unavailable, not observed, and not applicable.
Field presence
Reported presence across a representative sample, separated by source or record type.
Change handling
How fields, enumerations, and schema versions are communicated across deliveries.
Acceptation de l'échantillon
  1. Can the record be joined?Identity and keys fit the destination model.
  2. Can absence be interpreted?Missing states do not become false evidence.
  3. Can change be managed?Schema versions and delivery dates remain explicit.

Freshness and history

Choose what should follow dataset discovery.

Use a current snapshot for a baseline, available history for change analysis, or move the selected collection into Data Feeds for recurring scheduled delivery.

Une prise de vue

Établissez une ligne de base en temps réel.

Use a dated collection for one-time analysis, model preparation, market mapping, or a controlled initial load.

Inspect
Observation window and source mix
Plan
Full initial delivery
Récurrents Flux de données

Gardez la portée sélectionnée à jour.

Move to Flux de données to define a recurring schedule, schema, and destination for structured delivery.

Inspect
Cadence and change pattern
Plan
Scheduled delivery
Profondeur historique

Mettez le dernier record dans le contexte.

Where history is available, align the requested time window and grain with trend analysis, backtesting, or change detection.

Inspect
Available dates and continuity
Plan
Backfill plus future refreshes
01ObserveSource and capture window02ValidateSchema and quality checks03VersionDated delivery and manifest04DeliverFull or incremental handoff

Delivery design

Receive data in the shape your platform can operate.

Select JSON, CSV, or Parquet for your application, warehouse, or data lake. Clear file organization, manifests, schema versions, and refresh behavior keep each handoff predictable.

JSON

Préserver les objets en nid et enregistrer le contexte pour les flux de travail d'application ou axés sur les documents.

Nested records

VCS

Fournir une distribution de table familière pour les outils d'analyse, les entrepôts et les importations contrôlées.

Flat tables

Parquet

Supports le traitement par type et colonnal pour des charges de travail analytiques plus importantes et l'ingestion de la base de données.

Columnar files

Samples and pricing

Price the records and delivery you need.

Dataset pricing reflects the selected scope, history, refresh schedule, format, and delivery route. Start with a relevant collection and sample, then size the production data package.

Pratique de l'échantillon

Testez les dossiers avant de mesurer la livraison.

Exécutez des lignes représentatives et le dictionnaire de données à travers les joints, les filtres, les calculs et les contrôles de qualité utilisés par votre pile de production.

  • Representative source and entity mix
  • Sample record and data dictionary
  • Field presence and missing states
  • Snapshot date and available historical scope
  • Format and receiving environment
Request a dataset sample

FAQ

Data Marketplace questions.

See how catalog selection and samples connect to Data Feeds, Data API, Web Archive, and Managed Data.

Match my data brief

What is the Data Marketplace?

The Data Marketplace is a catalog of ready-to-use structured web datasets. Browse by domain and source, then inspect fields and representative sample records before choosing a prepared collection.

How do I find a relevant dataset?

Start with the entity and source, then narrow by markets, fields, historical scope, and format. Use the category pages to compare adjacent collections or send us the brief for help choosing.

Can I inspect the data before choosing a dataset?

Yes. Request representative records with the data dictionary and source context, then run the joins, filters, calculations, and quality checks your production stack will use.

Can a dataset be filtered to my scope?

Start with the countries, categories, entities, dates, sources, and fields you need. We will map that scope to the relevant collection and delivery structure.

What if I need data on a recurring schedule?

Use Data Marketplace to find and sample the right prepared collection. Choose Flux de données when the workload needs scheduled structured delivery with a defined cadence and destination.

Which delivery formats are available?

Choose JSON, CSV, or Parquet for the selected collection. Delivery can also define partitions, compression, file naming, manifests, and the destination used by your data stack.

How is Data Marketplace different from Data API?

Choose Data Marketplace to discover and sample a prepared collection. Choose Data API when your application requests source-specific structured results on demand, or Flux de données for recurring scheduled delivery.

When should I choose Managed Data instead?

Choose Managed Data when the required sources, schema, matching, quality rules, cadence, or destination need an operated program beyond the available catalog. WebScrapingAPI then manages collection, extraction, maintenance, monitoring, and delivery to your defined requirements.

Your dataset brief

Turn your data brief into a usable collection.

Tell us the domain, entities, markets, fields, and dates. We will connect the brief to a relevant prepared collection—or to Data Feeds when you need recurring delivery.