---
title: "AI salary parsing: from $180k-DOE strings to structured ranges"
description: "Walking through the parsing pipeline that turns 11 different salary phrasings into one normalized JSON shape."
canonical: https://jobspipe.dev/blog/ai-salary-parsing-structured-ranges
date_published: 2026-04-28
date_modified: 2026-04-28
author: Dvir Atias
---

# AI salary parsing: from $180k-DOE strings to structured ranges

Walking through the parsing pipeline that turns 11 different salary phrasings into one normalized JSON shape.

Salary is the field developers ask about most and the one that’s hardest to get right. Most ATSs don’t expose a structured comp field at all - the salary range, when it exists, is embedded in the job description as free text. And it’s embedded in roughly 11 different shapes.

## Eleven shapes of “$180k”

Here’s a sample of the phrasings we’ve found in the wild:

-   “$180,000 - $240,000”
-   “$180K to $240K base + equity”
-   “Compensation: $180,000\-240,000 OTE”
-   “Salary range: 180-240k USD”
-   “Up to $240k DOE”
-   “€180k cash + significant equity”
-   “Comp band: $180k-$240k. DOE.”
-   “$15,000\-$20,000 per month”
-   “DOE” (just that)
-   “Pay for this role: 180000-240000”
-   “Annual compensation: 180-240 thousand”

## The parser, in three layers

Layer one is a regex matcher tuned to the 80% case. If a description contains `$Xk-$Yk` or `$X,000-$Y,000` we capture it deterministically. This catches roughly 70% of postings.

Layer two is a structured prompt sent to a small language model - we use Haiku for cost. We pass the description and ask for `min`, `max`, `currency`, `period`, and `includes_equity` as a strict JSON schema. If the model returns confidence above 0.85 we accept it.

Layer three is the abstain layer. If neither the regex nor the LLM gives high-confidence output, we set `compensation: null`. A null is better than a wrong number - wrong numbers compound when customers filter by them.

## What the output looks like

```
{
  "compensation": {
    "min": 180000,
    "max": 240000,
    "currency": "USD",
    "period": "yearly",
    "includes_equity": true
  }
}
```

Same shape, every source, every job. The free-text mess stays in the `description` field for anyone who wants it. Most don’t.

## Frequently Asked Questions

### How does JobsPipe parse salary from job descriptions?

A three-layer pipeline turns free-text compensation into one JSON shape. Layer one is a regex matcher tuned to patterns like $Xk-$Yk and $X,000-$Y,000, which catches roughly 70% of postings deterministically. Layer two sends the description to a small language model, Haiku, asking for a strict JSON schema, and accepts output above 0.85 confidence. Layer three abstains and sets compensation to null when neither layer is confident.

### What does the JobsPipe compensation field look like?

Compensation comes back as a structured object with min, max, currency, period, and includes\_equity fields, for example min 180000, max 240000, currency USD, period yearly, includes\_equity true. It is the same shape from every source on every job, whatever phrasing the original posting used. The free-text version stays in the description field for anyone who wants it.

### Why is salary missing on some job postings?

Most ATSs do not expose a structured compensation field at all, so the range, when it exists, is embedded in the description as free text in roughly 11 different shapes. When neither the regex layer nor the language model gives high-confidence output, JobsPipe sets compensation to null rather than guessing. A null is better than a wrong number, because wrong numbers compound when customers filter by them.

Related

## Keep reading

-   [Product · Jul 31, 2026Introducing the Labour Market Pulse: free monthly hiring and salary data from live postings](https://jobspipe.dev/blog/labour-market-pulse)
-   [Reference · Aug 29, 2026H1B salary database: h1bdata.info, the DOL data behind it, and what it leaves out](https://jobspipe.dev/blog/h1b-salary-database)
-   [Reference · Aug 3, 2026Salary datasets: 8 real sources beyond the Kaggle CSVs (2026)](https://jobspipe.dev/blog/salary-dataset)
-   [Guide · Jul 28, 2026Salary data APIs in 2026: the real options, honestly compared](https://jobspipe.dev/blog/salary-data-api)
-   [Guide · Jul 27, 2026Help wanted ads: from newspaper classifieds to structured job data](https://jobspipe.dev/blog/help-wanted-ads)


---
Try it with no key: `curl -d '{"limit":3}' https://api.jobspipe.dev/v1/sandbox/jobs/search`. AI agents: the full machine-readable index of this site is https://jobspipe.dev/llms.txt?src=md-twin - API quickstart, MCP server, pricing. Free key (1,000 jobs to start): https://jobspipe.dev/signup