Architecture work starts with an inventory of what already exists
Aliases: content inventory · list before architecture · existing-content register
What it is
The first material of information architecture is not a navigation sketch. It is a content inventory: every existing object registered with location, type, owner, last update, entry points, and whether it is still referenced. Architecture organizes this collection, not an imagined ideal collection. Drawing a tree before the inventory draws a wish structure, which later splits on unregistered pages, PDFs, old campaigns, and shadow sites.
The inventory answers “what do we actually have.” The audit answers “should it still be here.” Count first, judge second. Mixing judgment into the first register will omit objects nobody wants to see.
Why it happens
Unregistered content still surfaces from search, inbound links, bookmarks, and internal forwards. If the architecture only covers “main pages” named in a workshop, those surfacing objects become orphans without a class, or are stuffed into the least-unlike bucket. The inventory turns hidden volume into computable input: type mix, staleness mix, ownerless mix. Those numbers decide how many layers the tree needs, whether facets have fields to eat, and which entrances actually point at duplicates.
Inferring content from an org chart or a competitor’s nav draws classes with no objects, and misses classes whose objects were never recalled. The inventory cuts that inversion.
Studying it
Treat the inventory as a method, and measure the cost of architecting without one.
- Paradigms: align three sources—a crawl, a CMS export, and URLs that appear in search logs but not in the sitemap—then compare an inventory-based tree with a workshop-based tree on tree tests and dead ends.
- Independent variables: inventory coverage (registered / actually reachable), inclusion of downloadable files and expired campaign pages.
- Dependent variables: post-launch orphan-page rate, search hits on unclassified objects, empty buckets whose live count is zero.
- Methodological note: walking down from the homepage by hand systematically misses deep pages and files. Logs and a full crawl are not optional extras. The inventory’s fields (type, owner, date) decide whether a later audit can run; a fieldless inventory is only a URL list.
Where it stops holding
A product starting from zero has almost no “existing content”; the inventory’s objects become competitors, business objects, and sources about to be migrated, not the local CMS. Personalized dynamic pages cannot be registered URL by URL; register types and templates, not every instance. Huge sites can only sample plus a priority queue (high traffic, high conversion, legally required), but the sampling plan must name what was deliberately left unregistered. Unsampled must not be treated as nonexistent.
Applying it
- Week one of an IA project produces only the inventory: URL or object id, type, source system, last update, entries. No navigation draft in parallel.
- Align three sources: CMS, crawler, search and inbound-link logs. Anything that appears in only one source is flagged on its own.
- Inventory fields only need to support later deduping and staleness calls. Do not turn week one into a perfect metadata programme.
- Verify by sampling twenty addresses reached from on-site search or inbound links. If any is missing from the inventory, the architecture is still working on a wish structure.
Related
- Within the group: G1.07.2 Audits expose duplicate, stale, and orphaned content · G1.07.3 Uninventoried content will break a new architecture
- Adjacent: G1.05 Metadata · G1.11 Architectural scalability and evolution · T3.03 Content freshness and maintenance
- Search terms:
content inventory·content audit·information architecture