How our job data is produced

This page describes how a job posting becomes data here: where it comes from, who asserted each value, and where we simply do not know. It is written to be checked. Every rule below matches the deployed code, not an intention.

What this covers

It describes the aggregated job index that feeds the job board, the public interfaces and the datasets. Two things are deliberately out of scope: ranking and search relevance, which are a product decision, and the coworking data, which comes from a separate source.

A note on availability, so this page does not promise more than is open: the per-field provenance discussed below belongs to the data model and is emitted by the second version of the jobs API. That version is not generally available yet. The rules still govern the data itself, including what the job board and the interfaces that are open today show.

Where the postings come from

There are two ways into the index. Aggregated postings are pulled from public feeds: employer applicant-tracking boards, job boards, and meta-aggregators. Direct postings are placed with us by employers themselves. The distinction is not bookkeeping, it decides what we are allowed to assert: with our own posting we spoke to the employer, with an aggregated one only to the source. Every posting also carries the kind of source it came from, which is how much distance sits between the posting and the company.

Kind of sourceWhat it means
atsThe employer applicant-tracking board. The shortest path: the posting sits where the company maintains it itself.
job_boardA job board. The posting was published there, so the fields come from the board rather than straight from the company.
aggregatorA meta-aggregator. The posting already passed through a collecting layer before it reached us, so at least one more rewrite.
governmentA government source. None is active right now, see the section on the Bundesagentur below.
unknownThe kind of source is not classified. We do not guess it.

The Bundesagentur für Arbeit was dropped in August 2026

On 21 August 2026 we removed the German Federal Employment Agency as a source, on legal grounds and not for data quality. Their responsible department confirmed to us in writing that the job-search interface is an internal interface for their own app and that reading postings out of it is not permitted. Separately, their terms grant usage rights in a posting only to the agency itself and to the cooperation partners the employer selected. We were never selected, so we never held a right to republish the posting text at all. That is why the source was removed rather than merely de-indexed.

That was an intervention in volume, not in quality. How large the index is today is further down, with an as-of date. We deliberately name no before-and-after delta here, because it would not be checkable on this page. What remains is mostly employer applicant-tracking boards and job boards whose terms actually cover redistribution. A later cooperation agreement could reopen the source, through the interface built for that purpose.

What we are not allowed to pass on

Some sources permit display on our own pages with attribution but forbid republication beyond it. For those we pass on neither the posting text nor the apply URL to third parties: both fields go out empty. One limitation belongs here, because it cuts against the rest of this page: a field withheld on policy grounds currently looks exactly like a field we never had. We emit no marker that tells the two apart, so an empty posting text from one of these sources must not be read as an employer who wrote nothing. Anyone who wants to read the posting can do so on its detail page here.

How work arrangement is decided

At ingest a rule-based classifier decides from structured signals. It is precise but blind to meaning: a "Remote Operations Producer" produces remote broadcasts and is not a remote job, and an on-site role whose body merely mentions the word remote looks identical to a pattern. So on top of it a small language model reads the full posting and returns remote, hybrid or onsite.

Where a model verdict exists it takes precedence over the rule tier. Only a verdict at or above the configured minimum confidence is acted on, high by default. An unsure verdict is discarded and the posting keeps its rule-based tier. The reason is uncomfortably concrete: a cheap misjudgement must not be able to make a genuine remote job disappear. If a source later edits the posting text, the verdict is formed again, detected through a checksum of the text the old verdict came from.

Remote and hybrid stay separate and are never folded into "remote-ish". An onsite verdict does not expire the posting: the row stays active, is filtered out when we serve, and its detail page answers 404. The difference matters, because an expiry date would be a statement about the job, while this is a statement about our classification. Where the arrangement cannot be determined at all, it reads unknown.

How a salary is resolved

Salaries are resolved when we serve, in three tiers, in this order: the figure stored at ingest, a fresh rule-based read of the current text, an extraction by a language model. If no tier finds anything, the field stays empty; the one exception is the Minijob legal derivation described below, which is not a stated salary but a statutory ceiling. When we resolve, we never splice a bound from one tier onto a bound from another: each tier is taken as a whole pair, because mixing them could invert a range, for instance at least 5000 with at most 4000. One case is honest to name, though. At ingest each bound is filled separately, so a feed can supply the lower bound while our own reading of the text supplies the upper one. That mixed pair is stored and later treated as a single tier, and we credit the whole range to our own extraction. That understates the source rather than putting a figure in its mouth that it never stated.

ProvenanceWhat it says
source_structuredThe source supplied the salary as a structured field. A third-party assertion the employer might dispute, but not one of ours.
nomado24_extractedWe read the figure out of the posting text, by rule or by model. Our reading, offered as ours, even when we are confident.
legal_derivationDerived from a named statutory rule. Exactly one case so far: for a German Minijob the statutory monthly earnings ceiling of 603 euro is also the salary ceiling, even when the posting states no figure.
unknownNo provenance can be established. This includes rows ingested before the provenance column shipped: that we no longer know which tier set the figure is no reason to credit it to the source.

The derived Minijob ceiling ships with its rule id (de.minijob.monthly_earnings_ceiling), version (2026-01) and effective date, so a consumer can re-derive it or throw it away. The ceiling is raised in most years, which is why the figure without a version would be worthless. It is explicitly not an employer offer: a Minijob may pay any amount up to it. All amounts are normalized to monthly euro; the period the posting itself stated, an hourly rate for instance, is carried separately.

The most important thing to say about salary is an absence: the large majority of postings state no salary at all. We do not fill that gap with estimates, industry averages or ranges from comparable roles.

How often each tier fires

This distribution is the actual evidence for everything above, which is why it sits here and not only in the coverage section. Measured on a sample of 2000 postings on 2026-08-25, over the served set. Unlike the figures further down it cannot be recomputed from a public address today, so the sample size and the date travel with it.

ShareWhat it counts
8.1%A salary is present at all, from whichever of the three tiers.
7.5%The posting itself states a figure, so either a structured source field or a figure we read out. This is the value for any sentence of the form "this many postings state a salary".
0.6%The source supplies the salary as a structured field. A very small share at the sample date, and that is the real finding; since the Personio mapping of 26 August 2026 it has been growing.
6.9%We read the figure out of the posting text. The large majority of all salary values.
0.15%The derived statutory Minijob ceiling. These postings state no salary and are therefore not part of the "the posting states a figure" row.
5.8%What the public statistics dataset reports. Lower, because it counts only the stored columns and not the re-read that happens when we serve.

So nearly every salary here is our reading rather than a statement by a source. That is exactly why each salary says which tier set it: without that label a structured source field and a figure we read out of prose look identical, although they are worth very different things.

How employment form is decided

We carry two vocabularies at once. One is the coarse schema.org set that search engines facet on: full time, part time, contractor, intern, temporary. The other is the finer set, which distinguishes Minijob, working student, apprenticeship, internship, freelance, contract, temporary, part time and full time. Since 26 August 2026 the structured statements of the source itself (Personio, for instance) flow in first, unioned with our own derivation. Provenance is recorded for the list as a whole, deliberately humbly: only when every form came from the source does it count as a source statement, otherwise as our classification.

The two are unioned rather than substituted; the finer forms come first (source-stated before our derived ones), then schema.org. The reason is the schema.org part-time bucket: it cannot tell a Minijob from a working-student role from ordinary part time. Replacing it with our form would mean anyone filtering on part time loses exactly those postings. Because both values are present, a part-time filter still matches them, and whoever means Minijob can say Minijob. In the other direction schema.org folds contract and freelance into one value, and we do not invent that distinction into a tier that never made it.

An empty list means not determined. It does not mean that no employment form applies.

Freshness and verification

Five timestamps that are easily confused. They deliberately say different things, and only one of them is a statement by the employer.

FieldWhat it means
publishedAtThe publication date as the source states it. Where the source states none, the field is absent. We do not substitute the date we found the posting.
firstSeenAtWhen we first saw the posting. An observation of ours, not a statement by the employer.
lastVerifiedAtWhen we last verified it, always together with the method: either we called the original URL and it answered, or the posting was still present in the feed. Both are real evidence, and the weaker one is never dressed up as the stronger one. A call that led nowhere, meaning a timeout, a server error, a rate limit or a bot wall, does not count as a fetch: the posting falls back to feed presence and to that timestamp. So where this field names the fetch, something really did answer.
updatedAtWhen the representation we publish last changed, for instance title, text, company or salary. A re-sighting on its own is not a change, otherwise the whole index would read "changed" every four hours.
expiresAtThe end of a rolling 30-day validity window from the last sighting, or an earlier end date where the source names one. So usually our window, not an application deadline set by the employer.

How a posting is re-checked

Broad aggregators often keep echoing a posting long after the role was filled at the employer. It then stays visible in the feed and would never go stale while its apply link is already dead. So we actively fetch the original URLs of active postings, least recently checked first, in batches, and never twice within twelve hours. Only 404 and 410 count as definitively gone and end the posting. Timeouts, server errors, rate limits and bot walls count as inconclusive, and inconclusive means keep. Deleting an open role because of a brief outage would be the more expensive mistake. That choice has a price, but it does not land on the timestamp above: an inconclusive probe keeps the posting alive without counting as a confirmation. We record what each check concluded, and only a fetch that actually answered is reported above as a fetch. Everything else falls back to feed presence. A posting behind a bot wall therefore does not look freshly confirmed. Postings without a fetchable original URL are left out of this check entirely, rather than counting as checked when no request ever went out.

Ingest runs every four hours, and each run refreshes sightings, checks links and ends what has expired. Four hours is the cadence of the job, though, not a publication guarantee. Every source also carries its own minimum fetch interval and is skipped until that has elapsed: 6 hours for Remotive, 24 hours for Adzuna. A posting that goes live just after its source was last read therefore waits for the next run that source is due in, which can be considerably longer than one window.

The heart of it: absence is published as a fact

The rule everything else follows from: missing beats guessed. Where we do not have a value the field stays empty, and the emptiness is the statement. Where we do have one, we say who asserted it. The data model does this per field, not per posting: a non-obvious field carries a provenance entry with its kind, source, observation time, method and method version. That lets a consumer decide for itself how much a value is worth instead of trusting an overall grade from us.

Kind of provenanceMeaning
employer_factThe employer stated it. Reserved for our own postings. The aggregated output does not currently use this kind, because nobody there spoke to us directly.
source_factThe source stated it. A third-party assertion the employer may still disagree with.
nomado24_classificationWe derived it, by rule, heuristic or model. Our reading, labelled as ours.
legal_derivationIt follows from a named statutory rule, which ships with its id, version and effective date.

One assurance we explicitly do not give: that a provenance can be established for every value. On a small share of postings, mostly ones ingested before the provenance column shipped, a salary is present and its provenance reads unknown. We then show both, the value and the admission. The alternatives would be to suppress the figure or to attribute it to a tier that may never have set it, and both are worse.

A provenance is never upgraded. A value we read out of prose stays our classification even when we are highly confident. Confidence is not provenance, and conflating the two is exactly the trick this page exists to prevent.

In practice that means a list of things we deliberately do not do:

  • We assign no occupational category, not even where a source states its own (Personio does): without a cross-source vocabulary any mapping would be invented. An empty category list is more honest.
  • We do not derive a region such as DACH or EEA from "Germany". That a role sits in Germany does not tell us the employer would accept an Austrian applicant. An empty region list means not determined, never not eligible.
  • We do not parse a city or a sub-national region out of the location hint, and for a multi-country scope we name no representative country.
  • We do not assert whether a salary is gross or net. German figures are quoted gross by convention, but no source we ingest states it machine-readably.
  • We do not ship a matched text snippet alongside a salary. We do not retain one, and reconstructing one would be fabricating evidence.
  • We do not publish a model summary whose underlying text has since changed. It would be a statement about a posting that no longer exists in that form.
  • We treat a company as a name, not as a verified legal entity. An employer publishing through two systems under two spellings appears twice, and we claim nothing else.

Coverage, with an as-of date

A methodology without numbers is a claim. The figures below describe how complete the index actually is, including the uncomfortable parts. The population is every posting active at the time of collection.

As of:

FigureValue
Active postings in the index3,588
with a stored salary figure (public dataset)5.8%
fully remote (ingest-time column)48.8%
hybrid (ingest-time column)51.2%
scoped to Germany57.8%
first seen in the last seven days14.9%

How to read these numbers

  • The figures come from the public statistics dataset and can be recomputed there. They are frozen here so a citation stays verifiable later. Current values live in the dataset itself, with its own collection timestamp.
  • The remote versus hybrid split is the column set at ingest, counted over the whole active table, so before the model verdict. On the job board the model verdict sits on top of it: it moves individual postings from remote to hybrid and removes onsite ones. The remote share we actually serve is therefore lower than the one given here.
  • Postings are counted, not employers. A company with twenty roles appears twenty times.
  • We deliberately publish no coverage rate for employment form. The public dataset computes none, and a number nobody can recompute does not belong on this page. Per posting, the value and its provenance travel with the record itself.
  • The salary share in this table is the public dataset’s, which counts only the stored columns. The full distribution across all three tiers is in the salary section and is higher, because we re-read the text when we serve. The derived Minijob ceiling counts as a stated salary in neither figure.

Dataset to recompute from: www.nomado24.de/remote-jobs-statistik.json

Statistics page with current figures

If something is wrong

If a posting carries a value that is not right, we want to hear it, with the address of the detail page. For an aggregated posting we cannot correct the source, but we can end the posting or fix our own classification. We would rather have classification errors reported than undiscovered: a methodology is only as good as its correction loop.

Contact: anton.petuchow@nomado24.de

Next