An incident page names a service. The catalog says owner: artist-relations-team. The on-call rotation for that team was dissolved six months ago. The page is stale, the rotation is empty, and the first person to notice is whoever is holding the pager when the alert fires.
This is not a Backstage bug. Backstage’s catalog is a registry of entities, not a reconciliation engine against your HR system. The descriptor format documents what owner is and where it lives; it does not promise that the string resolves to a living team. Treating the catalog as if it did is the mistake. The audit below is the correction: a scheduled, repeatable check that compares what catalog-info.yaml declares against what the catalog actually ingested, with a named owner and a published cadence.
What the descriptor format actually guarantees
Backstage’s descriptor format documentation is explicit about the shape of an entity. The root fields are apiVersion, kind, metadata, and spec. For a Component, owner is a field under spec:
spec:
type: website
lifecycle: production
owner: artist-relations-team
system: public-websites
That is the entire contract. owner is a string in a YAML file. The catalog stores it. Nothing in the descriptor format says the string must resolve to an existing Group entity at ingestion time, and the documentation does not claim the built-in validation rejects an unresolved owner. If you assumed otherwise, that assumption is the first thing to retire.
Two other documented facts matter for the audit:
- Entity names are unique per kind, within a given namespace, at any point in time, and the constraint is case insensitive. Names may be reused after an entity is deleted.
- When referring to an entity in a different namespace, you must use the
<namespace>/<name>syntax. In the default namespace, the namespace part can be omitted as a shorthand.
Those two rules are where most naive audits break. A script that compares bare owner strings against bare group names will flag correct cross-namespace references as broken, and will miss the case where a name was reused by a different entity after deletion.
Relations are an output, not an input
The descriptor format documentation states that relations is a read-only list, and that entity descriptor YAML files are not supposed to contain the field. Relations are derived by catalog processors from the entity data. A relation like { "type": "ownedBy", "targetRef": "group:default/dev.infra" } is computed, not authored.
This changes how you audit. You do not edit relations to fix ownership. You edit spec.owner in the descriptor, or you fix the ingestion of the Group entity, and the relation recomputes. If you find an ownedBy relation whose target group does not exist as an ingested entity, that is a signal about ingestion or about a stale reference, not about the YAML’s owner field. Treat relations as a derived output to diff against, never as the source of truth.
The four-pass audit
Run this on a schedule. The cadence is a decision for the platform team, but the audit itself should be a script or a query, not a manual review, so that the output is comparable across runs.
Pass 1: every Component has a non-empty spec.owner
Query the catalog for kind: Component and check that spec.owner is present and non-empty. This is the cheapest pass and it catches the entities that were scaffolded without an owner, or whose owner field was removed during a refactor. The catalog index page includes a default owner filter, so you can eyeball the empty case in the UI before scripting it, but the script is what you keep.
Pass 2: every owner string resolves to a Group entity
For each spec.owner value, resolve it against the set of ingested Group entities. Two rules apply:
- If the owner string contains a
/, treat it as<namespace>/<name>and resolve within that namespace. - If it does not, resolve within the default namespace.
A bare string comparison will produce false positives here. The documentation is clear that cross-namespace references require the qualified form, so an audit that ignores namespaces will report correct references as missing. Build the namespace handling into the script from the first run, not after the first false-positive review.
Pass 3: every referenced Group actually exists as an ingested entity
This is the pass that catches the failure mode from the opening. A service’s spec.owner points at a team that was merged into another team six months ago. The old Group entity was never removed from the catalog, so the name still resolves. Pass 2 passes. The org chart has moved on. The audit does not notice.
The fix is to make the org chart ingestion authoritative. Backstage’s external integrations documentation describes three ways to bring external data into the catalog: entity providers, custom processors, and incremental entity providers. An entity provider sits at the edge of the catalog as an original source of entities and gives you full control over when and how data is fetched. A processor runs inside the catalog’s processing loop and can enrich, validate, or transform entities after ingestion.
If your HR system or IdP already owns the org chart, ingest Group entities from it via an entity provider rather than hand-maintaining YAML. Then Pass 3 becomes: does the referenced group exist in the ingested set, and is it still present in the source system? A group that exists in the catalog but not in the source is a finding in its own right, and it is the finding that catches the merged-team case.
Pass 4: flag entities carrying backstage.io/orphan
The well-known annotations documentation states that backstage.io/orphan is injected by the catalog itself on entities found to have no registered locations or config locations keeping them active, and should never be added manually. An orphaned entity is one that the catalog is holding but no longer tracking from a source. For an ownership audit, an orphan is a candidate for removal or re-registration, and either way it needs a human decision. Flag it; do not auto-delete it.
The same page documents backstage.io/managed-by-location, which is added automatically by the catalog when it fetches data from a registered location and is not meant to normally be written by humans. That annotation is how you trace an entity back to the file that produced it, which is what you need when the audit output says “this owner is wrong” and someone has to open a PR.
Why the audit output belongs in a design doc
The output of the four passes is a table with four columns: entity ref, declared owner, resolved group ref, status. That table is the artifact. It is what you attach to a deprecation calendar entry when a team is being dissolved, and it is what you paste into a design doc when someone asks why the platform team is spending time on catalog hygiene.
The status column should distinguish at least four cases:
- Resolved — owner string resolves to an ingested group that is present in the source system.
- Unresolved — owner string does not resolve to any ingested group.
- Stale — owner string resolves to a group that exists in the catalog but not in the source system.
- Orphaned — entity carries
backstage.io/orphan.
“Stale” is the category most audits miss, and it is the one that produces the incident-page failure. A group that was deleted from the source system but never removed from the catalog will keep resolving. The audit has to compare against the source, not just against the catalog’s own contents.
The namespace trap, in detail
The descriptor format documentation gives a concrete example of why namespaces exist: importing users and groups from an HR system into the default namespace, while also ingesting users from a GitHub Enterprise installation that may have the same names. The GitHub ingestion goes into a separate namespace to avoid collisions.
For an ownership audit, this means an owner string like dev.infra and an owner string like ghe/dev.infra are different references. The first resolves in the default namespace; the second resolves in the ghe namespace. A script that strips the namespace or compares only the name part will conflate them. The documentation is explicit that when entities are in a different namespace, you need the <namespace>/<name> syntax, and that using namespaces typically means you need to explicitly specify the namespace when referring to the entity.
Build the resolution logic to match the documented syntax exactly. If the owner string contains a slash, split on the first slash and treat the left side as the namespace. If it does not, use the default namespace. Do not guess.
What the catalog does not do for you
Three things the catalog does not do, based on the documentation reviewed here:
- It does not validate that
spec.ownerresolves to an existingGroupentity at ingestion time. The descriptor format documents the field’s location and shape, not a referential integrity check. - It does not automatically remove entities when their owning team is deleted from an external system. The
backstage.io/orphanannotation exists precisely because the catalog holds entities that no longer have an active source, and it requires human action to remove them. - It does not reconcile against your org chart. That reconciliation is the audit you build, and it is the thing that has to run on a schedule.
The catalog backend, from a storage and API standpoint, does not care about the kind of entities it stores, per the extending the model documentation. That is a feature: it means you can model your org chart as Group entities without fighting the storage layer. It is also the reason the catalog will happily store an owner string that points at nothing.
Making the audit repeatable
The audit is only useful if it runs on the same cadence and produces comparable output. Three practices make that work:
Ingest the org chart from the system that owns it. Hand-maintained Group YAML drifts. The external integrations documentation describes entity providers as the most common integration pattern for syncing a remote system into the catalog on a schedule or in response to events like webhooks. Use that for the org chart, and the audit’s “stale” category becomes meaningful.
Export the catalog, do not scrape the UI. The catalog customization documentation describes a catalog export feature that exports the currently filtered catalog entities in CSV or JSON format. Filter by kind and owner, export, and diff against the previous run. That is a reproducible input for the audit script.
Name an owner for the audit itself. The catalog is an internal product. The audit is part of its SLA. If no one owns the audit, the findings accumulate and the table becomes another stale artifact. The platform team that owns the catalog owns the audit, and the deprecation calendar entry for a dissolving team is the trigger to re-run it.
What to do when the audit finds a mismatch
The finding is the mismatch between the declared owner and the resolved group. Which side is wrong depends on the case:
- Unresolved owner, group never existed: the descriptor is wrong. Open a PR against the
catalog-info.yamlto pointspec.ownerat the correct group. - Unresolved owner, group was deleted: the descriptor is stale. Same fix, but also check whether the deletion was intentional and whether other entities reference the same group.
- Stale group, owner string still resolves: the ingestion is wrong. The group should have been removed from the catalog when it was removed from the source. Fix the entity provider or the source sync, not the descriptor.
- Orphaned entity: the entity has no active source. Decide whether to re-register it or remove it, and record the decision.
Do not fix relations. They are derived. Fix the input, and let the catalog recompute.
FAQ
Does Backstage reject a catalog-info.yaml whose spec.owner does not resolve to a Group entity?
The descriptor format documentation reviewed here does not state that it does. It documents owner as a field under spec and documents the shape of the entity, but it does not document a referential integrity check against Group entities at ingestion time. Treat the resolution check as something you build, not something the catalog provides.
Can I audit ownership from the relations field instead of spec.owner?
No. The descriptor format documentation states that relations is a read-only list and that entity descriptor YAML files are not supposed to contain the field. Relations are derived by catalog processors. Use them as an output to diff against, not as the input to edit.
Why does my audit report missing owners that are actually correct?
The most common cause is namespace handling. The documentation states that cross-namespace references require the <namespace>/<name> syntax, and that in the default namespace the namespace part can be omitted. A script that compares bare names will flag correct qualified references as missing. Build the namespace resolution into the script from the first run.
How do I know if a Group entity is stale?
Compare the catalog’s ingested Group entities against the source system that owns the org chart. If a group exists in the catalog but not in the source, it is stale. This requires ingesting the org chart from the source via an entity provider or processor, per the external integrations documentation, rather than hand-maintaining Group YAML.
What does backstage.io/orphan mean for an ownership audit?
The well-known annotations documentation states that backstage.io/orphan is injected by the catalog on entities found to have no registered locations or config locations keeping them active, and should never be added manually. For an audit, an orphaned entity is one that needs a human decision: re-register it or remove it. Do not auto-delete it.
Should the audit run on a schedule or on demand?
On a schedule, with a named owner and a published cadence. An on-demand audit triggered by an incident is a cleanup project, not a control. The catalog is an internal product; the audit is part of its SLA. The deprecation calendar entry for a dissolving team is the trigger to re-run it, not the incident page.