If your team collects public web data, real-estate listings, business directories, agent profiles or tender notices, and feeds it into a pipeline that scores or ranks people, the rules you rely on are shifting. On 31 August 2026 the Attorney-General's Department released the exposure draft of the Privacy Amendment (Personal Data Protection) Bill 2026. Submissions close on 18 September 2026. The instinct that public means fair game is now the thing under review. Here is what the draft does, and a practical checklist to keep your collection and your decisions defensible.
The change in one paragraph
The exposure draft is the second tranche of Australia's privacy reform. It contains about 40 proposals that modernise how personal information is defined and handled. Three parts matter most if you scrape and pipe data. First, a new overarching fair and reasonable test is proposed to replace Australian Privacy Principles 3, 4 and 6. It weighs factors such as reasonable expectations, transparency, data minimisation and genuine choice. Second, the definition of personal information gets wider. Third, from 10 December 2026, entities must disclose the personal information used in substantially automated decisions. Submissions on the draft close on 18 September 2026, so the window to respond is short.
Why public is no longer the test
For years the working assumption was simple. If data sat on a public page, collecting it was safe. The draft reframes that. Scraping publicly accessible data that can be linked to an individual is still personal information, and it carries the same obligations as any other. The question is no longer whether the data was visible. The question is whether your collection is fair and reasonable.
That test weighs what a person would reasonably expect. Someone who lists a property or posts a professional profile expects it to be seen by buyers or clients. They may not expect it to be harvested, stored for years and fed into a model that ranks them. The gap between those two expectations is where risk now sits.
For your pipeline, this means the defence changes. That it was public is not enough on its own. You need to show why you collected the data, that you took only what you needed, and that a person could reasonably expect that use.
The numbers that show why regulators care
Regulators are not looking at this in the abstract. The volume of data incidents is rising, and the pipelines that move personal information are part of the picture.
- 1,205, highest since 2018
- OAIC data breach notifications, 2025
- Health service providers, 19%
- Top affected sector
- About 40
- Proposals in the exposure draft
- 10 December 2026
- Automated-decision disclosure duty starts
In 2025 the OAIC received 1,205 data breach notifications, the highest number since the Notifiable Data Breaches scheme began in 2018. Health service providers were the most affected sector at 19% of notifications, followed by financial services. Most breaches were attributed to malicious or criminal activity, with cyber hacking the primary cause. The draft also sets a tighter clock: once an entity has reasonable grounds to believe an eligible data breach has occurred, it must give the Commissioner a statement within 72 hours.
AI inferences count now
The wider definition of personal information is the part most teams miss. It captures AI-generated inferences, along with audio and video from wearables. If your model infers something about a person, that inference can be personal information in its own right.
Predicting a household's income band, scoring a lead's intent to sell, or flagging an applicant as high risk are all inferences about an identifiable person. You did not collect them from a public page. Your pipeline created them. Under the wider definition they still count, and the fair and reasonable test still applies.
This matters because inferences are often the most sensitive output of a pipeline, and the least documented. Raw fields are easy to point to. The logic that turns them into a score is usually buried in code, which makes it the hardest part to explain if someone asks.
A practical checklist for your pipeline
You do not need to stop collecting public data. You need to be able to justify what you hold and what you do with it. Work through this list one source at a time.
- Know what you collect and why. Map every field your pipeline pulls from each source, and write down the purpose. If you cannot state a purpose, stop collecting it.
- Minimise personal fields. Strip names, contact details and identifiers you do not need. The less you hold, the smaller your risk.
- Set a retention period. Decide how long each source stays in the pipeline, and delete on schedule rather than keeping data forever by default.
- Document your automated decisions. Record what personal information feeds each substantially automated decision, and the nature of the decision made.
- Update your privacy policy before 10 December 2026. The disclosure duty for automated decisions starts then, and OAIC guidance is expected by September 2026.
- Keep a lawful-basis note per source. For each site you scrape, record why the collection is fair and reasonable and what a person would reasonably expect.
What we would do about it
What to do this fortnight
You have two deadlines to work with. Submissions on the exposure draft close on 18 September 2026, so if the reform affects your work, make your view known before then. The automated-decision disclosure duty starts on 10 December 2026, which gives you a fixed date to aim your privacy policy and documentation at.
Start with one pipeline. Map what it collects, cut what it does not need, and write down why the rest is fair and reasonable. If you want a hand, book a pipeline review on the Australia Agents page and we will work through it with you.
Sources
- Privacy reform consultation, Attorney-General's Department (accessed 2026-09-07)
- Exposure Draft, Privacy Amendment (Personal Data Protection) Bill 2026 (accessed 2026-09-07)
- Notifiable Data Breaches statistics dashboard, OAIC (accessed 2026-09-07)