Australia, One Record at a Time · Season One, Article 12
Enter an old address into a browser and you may receive a 404 error, a domain-for-sale page or an entirely unrelated site. A web page usually disappears without ceremony. An organisation redesigns its site, a government department changes name, an individual stops paying for hosting, a server closes, a content-management system is replaced, or a platform deletes an account. A page that could be cited yesterday can look as though it never existed today.
The National Library of Australia’s Australian Web Archive resists this form of forgetting. Through Trove, readers can find historical snapshots of government sites, news pages, associations, personal blogs and sites created around public events. It turns browsing into archival research while making one difficulty unmistakable: preserving the digital past is not a matter of pressing “save” on the internet.
A web page is not a stationary document. It is a changing set of relationships. Code calls images, scripts request databases, links point elsewhere, interfaces rely on browsers, and access depends on accounts and platforms. To preserve one moment is to decide which of those relationships can and should travel with it.
What does a web snapshot preserve?
A web archive generally sends a crawler to a page, downloads accessible HTML, images, style sheets and other resources, and records the capture time. When the page is replayed later, the archive tries to assemble those saved components. The reader is not contacting the original server. The apparent page is a reconstruction produced by preserved resources in a new system.
“The 2014 website” is therefore not a simple object. Its home page may have been captured on one date, some internal pages on another, and images from an earlier cache, while external links were never archived. Dynamic menus, search boxes, maps, video, payment functions and logged-in areas may fail. A page can appear visually complete without reproducing its original interactions.
A snapshot is also not a screenshot. A screenshot preserves a visible surface but cannot provide searchable text or working links. A web archive tries to retain structure and navigation but depends more heavily on software, resource paths and replay technology. Neither is “the website itself”. Each preserves selected properties.
When using an archived page, its date should be understood as the capture date, not necessarily the date of authorship or last update. An undated policy statement could already have existed for months when collected. A careful citation records the archived address, capture time, page title and original URL, and compares multiple captures when necessary. One visit by a crawler is not a complete history.
Australia does not preserve the whole web in one way
The Australian Web Archive brings together three forms of collecting with different scopes. PANDORA has selectively preserved sites judged to have enduring value since 1996; curators can identify material around a community, subject or major event and work with publishers. The Australian Government Web Archive collects Commonwealth government sites more frequently. The annual Australian Domain Harvest gathers accessible .au sites at large scale and, according to the National Library, covers more than eighty per cent of the Australian web domain.
Selective collecting permits closer judgment and quality control and can identify important sites beyond mainstream attention. It will inevitably miss material a curator did not encounter. A domain-scale harvest offers breadth but is limited by time, storage, crawler restrictions, technical barriers and the definition of the national web. It does not mean every page and function of every site has been captured.
The methods do not merely duplicate one another. Different relationships compensate for different weaknesses. A community site may receive detailed treatment in PANDORA. A government policy may be collected several times in a year. Large numbers of small sites that nobody assessed in advance may retain at least a trace through the domain harvest.
Completeness is not achieved by saying “collect everything”. No feasible crawl can cover every website, every update and every function at once. A more trustworthy archive makes collection scope, frequency, failure conditions and selection principles visible, so readers can recognise where silence may have been produced.
Legal permission does not guarantee technical preservation
In 2016, Australia’s legal deposit provisions were extended to electronic publications, giving the National Library a clearer basis for collecting Australian online material. Web archiving involves making copyright copies; without legal authority, permissions and access rules, preservation itself can be restricted.
Legal entitlement does not automatically bring content into an archive. A site can block crawlers. Resources may sit on an overseas platform. A script may require a live connection to an interface that later closes. Audio and video may use restricted formats. Pages may be visible only after login. Social media is especially difficult: it is vast and continuously changing, its meaning often depends on account relationships and algorithmic ordering, and platforms frequently alter their terms and technical interfaces. The Library notes that most social-media content is not archived, although some Twitter material has been collected.
The resulting bias matters. Public, technically simple and crawlable websites are easier to retain. Conversation inside closed platforms, temporary stories, comment threads, private groups and algorithmically ordered feeds are more likely to vanish. A future researcher may see an institution’s official announcement without being able to recover how people responded to it on a platform.
The gaps in a digital archive are not random. Commercial platforms, permissions, web design, copyright and institutional resources shape them together. When a form of life increasingly occurs on an unarchivable platform, it also becomes less visible in future public memory.
A surviving domain may no longer contain the same page
Absence in a paper archive is often expressed by a missing file. Online absence can masquerade as presence. The same address may be acquired by a new owner. A page can be silently updated, a news headline changed, an old government guide overwritten, or an expired personal domain purchased for advertising or fraud.
This is why a web citation cannot rely only on “accessed on” a date. The live page proves what the present server supplied; it does not prove what the page showed last year. Archive captures give web evidence versions and time, making change itself available for study.
Comparisons can reveal when policy language changed, a promise disappeared, an agency reorganised its categories, or a crisis website gradually added information. They also require caution. If material is absent from one capture, the publisher may have removed it—or the crawler may have failed to retrieve a script or internal page. A difference is a lead for investigation, not always proof of an event.
Web archives do more than repair dead links. They allow us to follow how a state revises its public account of itself. Design, navigation hierarchy and linking also show what an institution considered important and what it wanted visitors to encounter first.
A searchable past is still organised by present systems
The Australian Web Archive can be searched through Trove by URL, keyword and subject. Search brings billions of captured resources within reach of ordinary readers, but adds another layer of mediation. How titles are extracted, full text indexed, duplicate pages handled and results ranked helps determine which past is discovered first.
A record that cannot be found has not ceased to exist, yet it may scarcely exist in practice. Researchers can mistake “no result” for “no event”, overlooking changes in spelling, old domains, text embedded in images, incomplete captures or indexing delays. Working with a web archive means moving beyond one keyword: following historical links, agency names, date ranges and related collections repeatedly.
Artificial intelligence will alter this relationship again. Summaries and question-answering systems can rapidly examine large collections, but a model may merge captures from different dates into one stable fact or erase differences in page version, capture quality and provenance. If the answer retains only a conclusion with no path back to a snapshot, the web archive has again been compressed into an output that cannot be inspected.
Intelligent tools for the digital past should preserve time, original URL, archive URL, capture status and the relevant evidence, and should let the reader return to page context. Speed can help discover material. It cannot replace the version identity of that material.
Preserving the web means migrating it continuously
Digital objects are often imagined as bits that never wear out. Bits can be copied exactly, but the software and environment that make them intelligible continue to change. Old formats lose support, character encodings become obsolete, browsers stop running plugins, domain structures change, storage media age and security requirements prevent old components from operating online.
Web archiving does not end when a crawl is stored. Institutions verify data, retain metadata, replicate backups, migrate storage, maintain replay software, repair search and adapt to new browsers without falsely rewriting the captured page. A snapshot opens today because people have continued to maintain a technical and institutional arrangement long after capture.
This maintenance attracts less attention than the discovery of a rare manuscript. It depends on budgets, standards, engineers, librarians, copyright policy and long-term organisational commitment. If one part ceases, preserved data can remain physically present while becoming inaccessible inventory.
Digital preservation therefore contains a paradox. To keep an object “the same”, custodians must keep changing the system that carries it. Continuity does not mean technological stasis. It means sustaining provenance, content, time and the boundary of alteration through each migration.
Who may be forgotten, and who must remain on the record?
The fact that a page was public does not mean all its contents should remain universally visible forever. A person may regret something written when young. Community material may be culturally sensitive. A false accusation can survive in an old capture after correction on the live site. Searchable names may continue harming victims.
At the same time, powerful institutions may prefer that old promises or failures disappear. An overly permissive removal system damages accountability, while absolute preservation can neglect privacy, safety and First Nations cultural rights. Web archives need processes for takedown, restricted access and legal duties, as well as records sufficient to explain why material is unavailable.
No single rule resolves every object. A defensible process distinguishes records of public power from vulnerable personal information, reduced search exposure from destruction, and temporary restriction from permanent removal. It also permits review. Remembering and forgetting are both institutional actions; neither should become invisible.
This series does not end with a complete archive
This season began with immigration forms, service files, census categories and protest documents, then moved through palace letters, promotional film, Hansard, a disaster inquiry, an automated welfare scheme and archived web pages. Every record disclosed something and excluded other experience according to its purpose, categories and technology.
An archive is not the nation’s completed autobiography. It resembles the residue of many actions. Government records in order to administer. An individual fills a form in order to return home. Parliament edits in order to make proceedings public. An inquiry assembles testimony to attribute responsibility. A library repeatedly migrates data so that future readers retain access. Different forces cause material to survive; survival is no guarantee of fairness or completeness.
Reading a primary record therefore cannot stop at “What does it say?” We also need to ask: Who made it, within which relationships? Which lives were converted into fields, numbers or images? How did it acquire authority? Who could correct it? Which technology keeps it legible? Where should we look for people who never entered the record?
The Australian Web Archive brings these questions into the present. Our era does not automatically leave more history merely because it produces enormous quantities of data. The more data grows, platforms close and pages change, the more survival and interpretation depend on continuing choices.
A country’s digital past will not wait naturally for the future. It must be noticed before pages disappear, migrated before technologies fail, and supplied with time and provenance before meaning is flattened. What an archive preserves is never content alone. It also preserves the relationships that allow content to remain understandable.
Primary record and further sources
- National Library of Australia: scope, collections and limitations of the Australian Web Archive
- Trove: search the Australian Web Archive
- PANDORA: Australia’s selective web archive
Continue reading: All articles in Australia, One Record at a Time
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.