Aller au contenu

Jobs & talent data

Job posting data with lifecycle, context, and limits intact.

Collect documented Google Jobs search and listing observations, or scope a maintained public-jobs feed that separates employers, postings, roles, locations, compensation, source context, and lifecycle events.

  1. 01
    Define the object

    Employer, posting, role, location, compensation, or event.

  2. 02
    Lock the source context

    Query, geography, language, source, and observation time.

  3. 03
    Agree lifecycle meaning

    Observed, changed, missing, inactive, and unknown.

Jobs record model

One listing page contains several objects—and none proves a hire.

Keep stable entities apart from source-specific posting text and time-bound observations. The model becomes useful for search, comparison, history, and analysis without overstating what a public page means.

01 · Primary objectDocumenté

Job posting

A source-specific public listing with its own identifier, URL, title, description, and observation state.

  • posting_id & URL
  • title & description
  • source dates & state
02 · EntityDocumenté

Employer

The public employer identity shown by the source, retained separately from any normalized organization.

  • source name
  • employer link
  • organization cues
03 · ClassificationLe pilote d'abord

Role

Normalized function, seniority, skills, and employment type derived under agreed rules.

  • function & level
  • skills
  • employment type
04 · ContextDocumenté

Location & work mode

Source text and structured cues for place, remote, hybrid, or on-site context.

  • location text
  • remote cue
  • normalized place
05 · ObservationLe pilote d'abord

Compensation

Displayed salary or rate, currency, period, range, and provenance—not an inferred market benchmark by default.

  • min & max
  • currency & period
  • source or derived
06 · EventLe pilote d'abord

Lifecycle event

A first-seen, content-change, last-seen, removal, or reappearance observation tied to one source.

  • event type
  • event time
  • source evidence
Interpretation boundaryNot assumed

A job posting is not a verified open vacancy.

A source can expose stale, duplicated, syndicated, removed, or reposted content. Public-page observations do not confirm headcount approval, applicant status, a filled role, or a hiring outcome.

Coverage brief

Know which jobs surface, input, and lifecycle you are buying.

Search results, individual listings, employer career pages, and normalized board feeds are different products. Source and field support is confirmed against representative public pages.

Source or page familyLes donnéesDelivered objectÉtat
Google Jobs APIQuery + locationJob-search result observationsDocumenté
Google Jobs Listing APIListing identifierIndividual public listing detailsDocumenté
Eligible public job pagesURL or source briefSource-specific posting recordLe pilote d'abord
Normalized feeds, enrichment & historySources + rule setLinked and versioned recordsLe pilote d'abord
Applicants, resumes & private ATSPrivate people dataExcludedPas standard

Posting schema

Keep source facts separate from normalized interpretation.

A trustworthy jobs record shows what the source said, what a rule normalized, when it was observed, and whether the page was present, changed, or unavailable.

01 · Posting

Source identity & content

posting_id
Source or delivery identifier for the listing.
title_raw
Job title exactly as exposed by the source.
description
Public listing text where the endpoint or scope supports it.
02 · Employer & role

Entity and classification

employer_raw
Employer name shown by the source.
role_normalized
Contracted function or title class, when included.
employment_type
Source value and normalized state where supported.
03 · Context

Place, work mode & compensation

location_raw
Location text shown with the posting.
work_mode
Remote, hybrid, on-site, unknown, or source-defined.
compensation
Value, range, currency, period, and derivation state.
04 · Lifecycle

Observation & change state

observed_at
Time the public source produced the observation.
first_seen / last_seen
History bounds within the contracted monitoring program.
collection_state
Observed, changed, missing, failed, or review.
Illustrative job posting record—not customer dataJSON
{
  "posting_id": "job_042",
  "source": {
    "listing_id": "source_818",
    "url": "https://jobs.example/818"
  },
  "employer": { "name_raw": "Example Labs" },
  "role": {
    "title_raw": "Data Engineer",
    "work_mode": "hybrid"
  },
  "location": { "text_raw": "Bucharest" },
  "compensation": { "state": "source_absent" },
  "lifecycle": {
    "first_seen": "YYYY-MM-DDThh:mm:ssZ",
    "last_seen": "YYYY-MM-DDThh:mm:ssZ",
    "state": "observed"
  },
  "schema_version": "jobs.v1"
}

Identity & matching

Preserve every posting before linking duplicates or reposts.

Source identifiers stay intact. Employer resolution, role normalization, and cross-source posting relationships are separate, explainable rule sets.

  1. 01

    Source key

    Listing ID + URL

    Retained with its source and observation time.
  2. 02

    Employer cues

    Name + domain + location

    Used only under the agreed entity rule.
  3. 03

    Posting cues

    Title + text + place + time

    Compared for exact, candidate, or repost relations.
  4. 04

    Relationship

    Source-only · linked · review

    No candidate is silently collapsed.

Freshness & lifecycle

A current snapshot and a posting history answer different questions.

Keep source-published dates apart from observation times. For recurring delivery, record change and disappearance without guessing the hiring outcome.

  1. Publishedpublished_at

    The date claimed by the source, when exposed.

  2. Observéobserved_at

    The page or result was collected in the requested context.

  3. Changedchanged_at

    A contracted field or content signature changed.

  4. Lifecyclefirst_seen / last_seen

    The posting entered and most recently appeared in tracked history.

Quality & missingness

Make absence and uncertainty usable.

Do not convert a missing salary to zero, an unspecified work mode to on-site, or an unavailable page to a filled role.

Observé

Required posting structure present

The public record passed the contracted source and field checks.

Changed

Material source fields changed

Prior and current observations retain a change reason or field diff.

Révision

Identity or normalization is ambiguous

The candidate remains inspectable rather than silently accepted.

Unavailable

No supported observation

Source absent, removed, access failed, and out-of-scope remain different states.

Structural checks

Required identifiers, URL shape, source state, timestamps, field types, and schema version.

Semantic checks

Compensation units, work mode, geography, employer cues, description presence, and lifecycle transitions where contracted.

Customer acceptance

The customer approves source eligibility, taxonomy, match rules, completeness thresholds, and decision use.

Operating model

Choose who owns the public-jobs data operation.

Compare the same posting lifecycle across infrastructure, documented APIs, scheduled feeds, and a managed program.

Maximum control

Build your collectors and posting model on proxy infrastructure.

Your team owns source discovery, extraction, schema, normalization, lifecycle rules, monitoring, maintenance, storage, and interpretation.
ResponsabilitéPropriétaire
Source, object, and field briefCustomer
Network access infrastructureWSA
Extraction and source schemaCustomer
Identity, lifecycle, and maintenanceCustomer
Storage, interpretation, and decisionsCustomer

Applications

Use one governed posting foundation across talent workflows.

The data supports market observation and product experiences; customers own analysis, employment-law obligations, and every decision involving people.

01

Labor-market research

Observe posting volume, roles, skills, employers, and locations across a defined public source set.

posting · role · employer · place · observed_at
02

Skills demand analysis

Map source-exposed requirements into an approved taxonomy with raw text retained.

description · skill cue · taxonomy · rule version
03

Compensation observation

Compare source-displayed ranges with currency, period, location, and missingness attached.

compensation · unit · place · provenance
04

Job search products

Power discovery and alert workflows from documented request-time records or a scoped feed.

query · result · posting · source context
05

Employer hiring signals

Track public posting observations over time without equating them to verified headcount or hires.

employer · posting · first_seen · last_seen
06

Location & remote trends

Study source-exposed work mode and geography with normalization state kept visible.

location_raw · normalized place · work mode · time

Representative pilot

Test the posting lifecycle before scaling the source list.

Use representative searches and listings to verify field presence, employer and role rules, compensation missingness, duplicate states, cadence, and downstream fit.

Scope a jobs data pilot
  1. 01

    Define

    Name queries, sources, geographies, objects, required fields, cadence, and excluded people data.

  2. 02

    Sample

    Include duplicates, reposts, multi-location roles, missing salaries, removed pages, and changed content.

  3. 03

    Inspect

    Review raw and normalized fields, match evidence, lifecycle events, timestamps, and quality states.

  4. 04

    Accept

    Agree schema, taxonomies, identity rules, cadence, quality thresholds, ownership, and delivery.

Evaluation FAQ

Resolve what one public posting can—and cannot—tell you.

These answers separate source observations from vacancy claims, people data, and employment decisions.

What jobs data can WebScrapingAPI provide?

The documented Google Jobs API supports job-search result observations, and the Google Jobs Listing API supports individual public listing details. Eligible public job pages can also be evaluated. Normalized multi-board feeds, compensation enrichment, company enrichment, and posting history begin with a representative pilot.

Does a job posting prove that a role is still open?

No. A job posting is not a verified open vacancy. It is a public source observation at a stated time. A source can be stale, duplicated, removed, reposted, or changed, so vacancy interpretation belongs to the customer.

What is the difference between a job posting and a role?

A posting is a source-specific published object with its own URL, identifier, text, and observed lifecycle. A role is a normalized concept such as title, function, seniority, or skills. Multiple postings may describe the same hiring need, and that relationship is only asserted when agreed matching rules support it.

Can jobs be tracked from first seen to closed?

Lifecycle history is pilot first. A recurring program can retain first_seen, last_seen, content-change events, removal observations, and a derived inactive state. A removed page is evidence that the source no longer exposed the posting, not proof that the employer filled or cancelled the role.

Can compensation and company data be normalized?

Source-exposed compensation, employer names, locations, and work-mode text can be retained. Cross-source normalization, currency or period conversion, salary inference, employer resolution, and company enrichment are pilot-first rules with explicit provenance and missing states.

Does this include applicant data or resumes?

No. Applicant data, resumes, private ATS records, candidate profiles, hiring-manager notes, and login-only recruiting systems are not standard scope. This page concerns eligible public job-posting observations, not people data.

Can the data be used for employment screening decisions?

WebScrapingAPI does not provide FCRA or employment-screening decisions, candidate eligibility judgments, or hiring recommendations. Customers remain responsible for lawful use, validation, governance, and decisions made with public jobs data.

How are duplicates and reposts handled?

Source identifiers and URLs are retained first. A pilot can evaluate employer, title, location, description, and time cues to label exact, candidate, reposted, ambiguous, or source-only relationships. Similar postings are not automatically collapsed into one vacancy.

What does a missing salary, location, or closing date mean?

The schema distinguishes a field not shown by the source, a value outside the requested object, a page unavailable in the request context, a collection failure, and a normalization needing review. Missing data is not silently converted to zero, remote, or closed.

Who maintains collection when a job source changes?

Proxy customers maintain their own collectors, parsers, and monitoring. WebScrapingAPI maintains documented API behavior within the product boundary. Scheduled and managed programs can include contracted extraction, normalization, quality monitoring, source-change maintenance, lifecycle history, and delivery.

Jobs & talent data

Start with a documented Jobs API or a representative source brief.

Bring the searches, listings, fields, geographies, lifecycle rules, cadence, and delivery your workflow needs.