---
title: "Indeed scraper in Python: how it works, what breaks"
description: "An Indeed scraper loads Indeed's search pages and reads the job results embedded in each page as JSON. Parsing takes a few lines of Python."
canonical: https://jobspipe.dev/blog/indeed-scraper
date_published: 2026-10-06
date_modified: 2026-10-06
author: Dvir Atias
---

# Indeed scraper in Python: how Indeed scraping works, where it breaks, and what hosted scrapers cost

An Indeed scraper loads Indeed's search pages and reads the job results embedded in each page as JSON. Parsing takes a few lines of Python. Fetching is the hard part: a plain request gets a challenge page, so a daily scraper needs a real browser, a pool of addresses and upkeep. Hosted scrapers and jobs APIs trade that work for a price.

Working Python for both halves of an Indeed scraper, the parser and the fetch, with what the fetch returned when we ran it. Then what breaks once it runs every day, what hosted Indeed scrapers charge, and the route that skips scraping.

Most Indeed scraper tutorials show the parser and stop. The parser is the easy half. The half that decides whether a scraper survives is the fetch, and it fails on the first run. The examples below are for developers who need Indeed postings as data inside a product or an analysis.

## How does an Indeed scraper work?

Start with the search URL. A search on Indeed is a GET request to `https://www.indeed.com/jobs` with a handful of query parameters:

-   `q` is the search text, for example a job title.
-   `l` is the location, such as a city and state.
-   `sort=date` orders results newest first, which makes paging stable from one run to the next.
-   `fromage` limits results to postings from the last N days.
-   `start` is the paging offset, moving in steps of ten.

The response is a full HTML page. Inside it, a script block assigns the page’s results to `window.mosaic.providerData["mosaic-provider-jobcards"]`. Within that object, under the key `mosaicProviderJobCardsModel`, sits a `results` array with one entry per job card, around fifteen on a page. Each entry carries the fields a scraper wants: `jobkey` (Indeed’s id for the job), `displayTitle`, `company`, `formattedLocation`, `pubDate` as a timestamp in milliseconds, and salary hints under `salarySnippet` and `extractedSalary` when the employer gave a range. The job itself lives at `/viewjob?jk=` followed by the job key.

So there are no HTML cards to parse with CSS selectors. Find the embedded object, decode it as JSON, and read the array.

Try it without a key

```
curl -d '{"limit":3}' https://api.jobspipe.dev/v1/sandbox/jobs/search
```

Only `/v1/sandbox/*` needs no key.

## A minimal Indeed scraper in Python

The parser first, because it is the part that works. It finds the model by name, decodes the JSON object that follows, and keeps one row per job key, since sponsored cards can repeat a job on the same page:

```
import json

ANCHOR = '"mosaicProviderJobCardsModel"'

def parse(html):
    at = html.find(ANCHOR)
    if at < 0:
        return []
    start = html.find("{", at)
    model, _ = json.JSONDecoder().raw_decode(html, start)
    jobs = {}
    for r in model.get("results", []):
        key = r.get("jobkey")
        if not key or key in jobs:
            continue
        jobs[key] = {
            "jobkey": key,
            "title": r.get("displayTitle") or r.get("title"),
            "company": r.get("company"),
            "location": r.get("formattedLocation"),
            "posted_ms": r.get("pubDate"),
            "url": f"https://www.indeed.com/viewjob?jk={key}",
        }
    return list(jobs.values())
```

`raw_decode` reads one JSON value from a position in a string and stops, which suits an object followed by more JavaScript. An empty list means the model was not on the page.

Now the fetch, in the form most tutorials show it:

```
import requests

def fetch(query, location, start=0):
    params = {"q": query, "l": location, "sort": "date", "start": start}
    resp = requests.get("https://www.indeed.com/jobs", params=params, timeout=20)
    print(resp.status_code)
    return resp.text

html = fetch("data engineer", "Austin, TX")
print(len(parse(html)), "jobs")
```

We ran this on October 6, 2026. It printed `403` and `0 jobs`. The body was a challenge page, with no job model in it.

## Why does a plain request get a 403?

Indeed’s search pages sit behind bot protection. A request that does not look like a browser is answered with a page that expects to run JavaScript before any results are shown, and an HTTP client cannot run it. Changing the user agent string does not change that. The usual next step is a real browser driven from code:

```
from playwright.sync_api import sync_playwright

def fetch_with_browser(url):
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=False)
        page = browser.new_page()
        page.goto(url, wait_until="domcontentloaded")
        html = page.content()
        browser.close()
    return html
```

Feed that HTML to `parse` and the job list fills in when the browser was served results. On a laptop, with a visible browser on a home connection, it often is. On a server it often is not, because requests from one data-centre address are challenged quickly. That gap between the laptop and the server is where most Indeed scraper projects stall.

## What breaks an Indeed scraper in production?

A scraper that runs every day meets five things a one-off script never does.

-   **The challenge changes.** When the protection in front of the site changes, every request fails at once.
-   **The embedded model is not a contract.** The key names above are Indeed’s internal names. A renamed field raises no error; it returns empty values until someone notices.
-   **Addresses wear out.** Residential addresses are billed by the gigabyte and a search page is heavy, so bandwidth becomes the main running cost.
-   **Closed jobs are invisible.** A search page shows what is listed now. Knowing that a posting has closed means fetching that job’s page again, for every job you hold.
-   **The same job arrives many times.** One posting turns up under several queries and locations, so deduplication by job key is the minimum, and it still misses a job advertised on other sites.

## What do hosted Indeed scrapers cost?

If you would sooner pay than maintain, there are three kinds of product, billed in different units.

**A scraper you rent by the result.** Apify is a platform of ready-made scrapers called Actors. The Indeed Scraper by misceres charges $5 per 1,000 listings on Starter, falling to $3 on Business, on top of the platform plan: paid plans from $19/month plus pay-as-you-go usage. You run and schedule it yourself. Our [Apify comparison](https://jobspipe.dev/alternatives/apify) covers when that is the right shape.

**A scraper API billed by the record.** Bright Data runs the scraper for you and returns records. The Web Scraper API, which includes ready-made LinkedIn and Indeed job scrapers, costs $1.50 per 1,000 records pay as you go, with 5,000 free records a month. Merging Indeed with other sources is still yours to build.

**An open-source library.** JobSpy is a free Python library with an Indeed scraper built in. The code costs nothing; the addresses it needs do, and the upkeep is yours. Our [JobSpy review](https://jobspipe.dev/blog/jobspy-review) covers what it returns and where it stops.

Whichever you pick, compare the price per posting you would keep, not per row returned: duplicates and reposts are billed like any other row.

## Is it legal to scrape Indeed?

That is two questions. One is the law, which depends on where you are, what you collect and what you do with it. The other is the site’s own terms. Read Indeed’s terms before you automate anything against the site, and take advice for a commercial product. Our guide to [whether web scraping is legal](https://jobspipe.dev/blog/is-web-scraping-legal) sets out the cases people cite and what they decide. This is not legal advice.

## How do you get Indeed jobs without a scraper?

Use a jobs API that already carries them. JobsPipe returns postings from 30+ sources in one schema, and Indeed is one of those sources, selected with a single filter:

```
curl https://api.jobspipe.dev/v1/jobs/search \
  -H "Authorization: Bearer jp_live_your_key_here" \
  -H "Content-Type: application/json" \
  -d '{
    "source_or": ["indeed"],
    "job_title_or": ["data engineer"],
    "job_country_code_or": ["US"],
    "posted_at_max_age_days": 7,
    "limit": 25
  }'
```

Each record has the same fields whichever source it came from: title, company, normalized location, salary in `min_annual_salary_usd` and `max_annual_salary_usd` when the posting states one, `date_posted`, an open or closed status, and a `url` back to the original listing. A job that appears on several sources comes back once. The free tier is 1,000 jobs to start.

A scraper is still the better route for a one-off research pull or a field no API exposes. For a product people rely on every day, the upkeep usually costs more than the plan. The [Indeed API guide](https://jobspipe.dev/blog/indeed-api-guide) has the build on the API side: pagination, deduplication and incremental sync.

## Sources

-   [Apify Store: Indeed Scraper by misceres](https://apify.com/misceres/indeed-scraper)
-   [Apify pricing](https://apify.com/pricing)
-   [Bright Data Web Scraper API pricing](https://brightdata.com/pricing/web-scraper)
-   [JobSpy on GitHub](https://github.com/speedyapply/JobSpy)

## Frequently Asked Questions

### Can you scrape Indeed with Python?

Yes, in two steps. Parsing is short: each Indeed search page embeds its results as JSON under mosaicProviderJobCardsModel, so you find that object, decode it and read the results array. Fetching is the hard step. A plain requests call gets a challenge page, so a repeatable scraper needs a real browser driven from code and, on a server, a pool of residential addresses.

### Why does my Indeed scraper return 403?

Because Indeed's search pages sit behind bot protection. A request that does not look like a browser gets a page that expects to run JavaScript first, and an HTTP client such as requests cannot run it. Changing the user agent does not help. A browser driven with Playwright can pass on a home connection; a single data-centre address is usually challenged quickly.

### Is there an Indeed API I can use instead of a scraper?

Not from Indeed for reading job search results. Indeed's public job-search API is retired, and the APIs that remain are partner APIs for posting jobs, receiving applications and managing ads. Reading Indeed postings as data means a jobs API that carries them, a wrapper over Google's job results, or a scraper.

### How much does a hosted Indeed scraper cost?

It depends on the unit. The Indeed Scraper by misceres charges $5 per 1,000 listings on Starter, falling to $3 on Business, on top of the platform plan. A scraper API bills by the record: the Web Scraper API, which includes ready-made LinkedIn and Indeed job scrapers, costs $1.50 per 1,000 records pay as you go, with 5,000 free records a month. An open-source library such as JobSpy is free, and the addresses it needs are not. Compare the price per posting you would keep, since duplicates are billed like any other row.

Related

## Keep reading

-   [Comparison · Jan 4, 2026Glassdoor scraper: what it costs, what breaks, and the API alternative](https://jobspipe.dev/blog/glassdoor-scraper-vs-api)
-   [Guide · Sep 9, 2026Indeed ghost jobs: why Indeed scores cleaner than most boards and where its stale postings hide](https://jobspipe.dev/blog/indeed-ghost-jobs)
-   [Guide · Sep 9, 2026Is Indeed an Applicant Tracking System?](https://jobspipe.dev/blog/is-indeed-an-applicant-tracking-system)
-   [Guide · Jul 25, 2026What "job closed or expired" means on Indeed (and what to do about it)](https://jobspipe.dev/blog/job-closed-or-expired-indeed)
-   [Comparison · Jul 13, 2026JobSpy GitHub review: the Python job scraper tested, and where it breaks](https://jobspipe.dev/blog/jobspy-review)


---
Try it with no key: `curl -d '{"limit":3}' https://api.jobspipe.dev/v1/sandbox/jobs/search`. AI agents: the full machine-readable index of this site is https://jobspipe.dev/llms.txt?src=md-twin - API quickstart, MCP server, pricing. Free key (1,000 jobs to start): https://jobspipe.dev/signup