PERFMETRIX — in data we trust
All articles
First-Party Data8 September 2026·9 min read

First-party data: the term everyone sells and nobody defines

Data a customer gave you directly, that you have a lawful basis to use for this purpose. Vendors quote the first half of that and skip the second.

DeBy Denis · Perfmetrix
On this page
  1. The definition has a sourcing half and a purpose half
  2. Which party collected it says nothing about how sensitive it is
  3. Data collected under consent can't be quietly repurposed
  4. The consent state is the join between your database and the bidding platform
  5. What it fixes about attribution is narrower than the pitch
  6. Four questions that test whether yours is usable

Short answer: first-party data is information a customer shared directly with you, which you hold a lawful basis to use for the specific purpose you're about to use it for. Both halves are load-bearing. Google's own customer data policies define it as "customer information that is shared directly with you by the customer" — that's the sourcing half, and it's the half every vendor deck quotes. The purpose half comes from data protection law, and it's why an email address collected for delivery updates isn't automatically an email address you may hash and send to an advertising platform.

Get the second half wrong and the first half doesn't save you. "We collected it ourselves" is a statement about provenance, not permission.

The definition has a sourcing half and a purpose half

Two different rulebooks are doing the work here, and conflating them is what makes the term slippery.

The sourcing rule comes from the platform. Google's customer data policies, which govern enhanced conversions, shop sales uploads and Google-engaged audiences, say you may only upload first-party data, "which is defined as customer information that is shared directly with you by the customer", gathered from your websites, apps, physical shops or other places where customers interact with your business. Their examples are concrete: people who bought something, registered for marketing messages, requested a quote, or signed up for an account or loyalty scheme. There's a carve-out for manufacturers whose products sell through retailers, and it comes with conditions — written assurance that the retailer complies, and written agreements where the law requires them.

The purpose rule comes from UK GDPR. Article 5(1)(b) requires personal information to be "collected... for specified, explicit and legitimate purposes and not further processed... in a manner that is incompatible with the purposes for which the controller collected the data". The ICO's guidance on that principle — updated on 23 March 2026 to reflect the Data (Use and Access) Act — puts it as: be clear from the outset why you're collecting, document it, tell people, and make sure any reuse is compatible with the original purpose.

So a usable definition is: data a customer gave you directly, for a purpose that covers what you're now doing with it. Anything that satisfies only the first clause is data you possess, not data you may deploy.

Which party collected it says nothing about how sensitive it is

The party numbering is a description of the collection route, and it carries no legal privilege. A hashed email you collected on your own checkout page is personal data in exactly the way a hashed email bought from a broker is, and UK GDPR applies to both.

That surprises people, so it's worth pinning to the regulator's own words. On hashing and similar techniques, the ICO's position is that pseudonymisation "is effectively only a security measure. It does not change the status of the data as personal data", and Recital 26 treats pseudonymised information that could be attributed to a person with additional information as information about an identifiable person. Hashing improves security. It does not convert regulated data into unregulated data.

What first-party collection genuinely gives you is three practical advantages, and they're worth having:

  • A relationship to point at. You can explain to a person, plausibly, how you got their details.
  • Control of the record. You can honour a withdrawal, correct an error and answer an access request without asking a third party for help.
  • Durability. It doesn't depend on a third-party cookie, so browser policy changes don't quietly delete it.

None of those is a permission. They're the conditions that make obtaining a permission realistic.

This is the failure that shows up in audits, and it's the one that makes the whole cluster worth writing about.

The ICO is explicit that the rules on reuse are stricter where the original lawful basis was consent. For personal information collected under consent, a new use is compatible only if you get consent from the person for the new specified, legitimate and explicit use, or if the reuse falls into a narrow set of conditions — complying with a data protection principle, a purpose listed in annex 2 of the UK GDPR, or safeguarding a public interest objective in article 23(1)(c) to (j) where it isn't reasonable to expect you to get new consent. Ad measurement and audience matching aren't on that list.

The consent request itself has to have been specific enough to cover the new use in the first place. The ICO's guidance on obtaining consent requires that a request name "the names of any other controllers who will rely on the consent", and states plainly that "consent for categories of third-party controllers will not be specific enough". A tick box saying you may share data with "our advertising partners" doesn't describe an upload to a named platform.

Two everyday examples of the gap:

  • A newsletter opt-in taken on a competition entry page, later used as the source list for an audience upload. The person consented to a newsletter.
  • Delivery-address data collected to fulfil an order under contract, later hashed as match keys. Contract covers the delivery. It doesn't cover the matching.

Neither is exotic and neither is malicious. Both are what happens when the marketing team inherits a database and reads "first-party" as a clearance. Collected-by-you is not the same as consented-for-this-purpose, and only the second one decides what you may send.

Most of these problems start at collection, on pages nobody has audited in two years. Our cookie and tracker scanner loads your site in a real browser and shows which cookies and trackers fire before anyone clicks Accept, and which consent platform — if any — is actually enforcing.

Picture the chain, because every argument about first-party data is really an argument about one of its links.

  1. Collection. A person gives you an email address on your site, in a context that specified a purpose.
  2. Consent state. Your banner records what they agreed to, and the browser signals it. In Google's model the relevant signal is ad_user_data, defined in the gtag reference as consent "for sending user data to Google for advertising purposes".
  3. Storage. The identifier and the consent decision sit against the same record in your CRM or warehouse. If they live in separate systems that can't be joined, the chain is already broken.
  4. Transmission. You normalise and hash the identifier and send it. Google's Data Manager API takes email addresses, phone numbers and address components as SHA-256 hashes after normalisation, with a maximum of ten identifiers per person in a single event.
  5. Use. The platform matches, attributes and bids.

Step 3 is where most implementations fail, and it fails invisibly because every other step keeps working. The banner still shows, the tags still fire, the upload still returns success. What's missing is the ability to say, for any individual row you sent, which decision that person made and when.

Google's EU user consent policy makes that a contractual duty rather than a nicety for users in the EEA, the UK and Switzerland: you must obtain consent, retain records of the consent given, and provide clear instructions for revocation. A record you can't produce is a record you don't have.

What it fixes about attribution is narrower than the pitch

First-party data closes the gap for people you can identify. That's a real gap and closing it is worth money. It leaves two others open, and no amount of collection reaches them.

Anonymous traffic stays anonymous. Someone who browsed, didn't convert and left has given you nothing to match on. First-party strategy has no purchase on them at all — the honest tools for that population are modelling and aggregate measurement, not identity.

Denied consent stays denied. A person who declined is not a data-quality problem to be routed around. Where an identifier can't be sent, the correct system behaviour is not to send it, and the gap that leaves in reporting is the system working.

There's a third limit that's easy to miss: matching only helps if the platform recognises the person, which for Google means matching against signed-in accounts. Your ability to identify a customer and Google's ability to identify the same customer are different facts, and only one of them is under your control.

Where first-party data genuinely changes the picture is in feeding outcomes back — telling the platform what a lead was worth rather than that it happened. That's a different job from recovering lost hits, which is what server-side tagging actually recovers: a narrower list than the vendor pages suggest, and mostly about transport rather than permission.

Four questions that test whether yours is usable

Run these against a real table, not against the strategy deck.

  1. Where did each row come from? If the answer varies by row and nobody recorded which, the sourcing half of the definition is unproven for the whole table.
  2. What purpose was specified at collection? Find the actual wording that was on the page at the time. Not the current privacy policy — the one that was live when the row was created.
  3. Can you produce the consent decision for one named person, with a date? If it takes an engineer a week, you can't do it during a regulator's enquiry either.
  4. Can you honour a withdrawal all the way to the platform? Deleting a row in your CRM does nothing to an audience already uploaded. The removal has to reach the destination.

A "no" to any of these doesn't mean stopping. It means the next piece of work is in the collection layer rather than the activation layer — and that's the cheaper end of the problem to fix.

The measurement consequence of getting this right, and the plumbing between the consent record and the platform, is what we build. If your immediate question is the narrower one — how hashed identifiers rejoin a lead to a click, and what that can't do — that's enhanced conversions for leads.

Sources

  1. 1.Google Advertising Policies — Customer data policies (first-party data definition) · Checked 2026-09-08
  2. 2.ICO — Principle (b): Purpose limitation, updated 23 March 2026 for the Data (Use and Access) Act · Checked 2026-09-08
  3. 3.ICO — How should we obtain, record and manage consent? · Checked 2026-09-08
  4. 4.ICO — What is personal data? (pseudonymised data remains personal data) · Checked 2026-09-08
  5. 5.Google — Data Manager API, UserData reference · Checked 2026-09-08
  6. 6.Google — gtag.js consent reference (ad_user_data) · Checked 2026-09-08
  7. 7.Google — EU user consent policy · Checked 2026-09-08

Find out what your site leaks — in 30 seconds

Run the free consent checker on your own domain, or book a call and we'll walk your setup together.