BIBFRAME Hub family explorer

See how a book, film, opera or translation is connected to the other works in its bibliographic family: adaptations, sequels, parodies, translations, series.

The data is the Library of Congress BIBFRAME Hubs (a Hub is an authorized name/title heading with a URI, e.g. Austen, Jane. Pride and prejudice). About 918K Hubs are linked to at least one other Hub by 544K typed relations; that connected part of the graph is what you browse here. It loads from Parquet on Hugging Face and runs entirely in your browser, with no server.

How to use it: search for a title, pick a Hub, and the page (1) climbs from that Hub up to the original work, then (2) shows everything related to that work, most surprising first. Click any title to make it the new center. Background and sources are in Design notes at the end.

// Load the two Parquet files into an in-browser DuckDB (~97 MB on first load). We fetch the
// bytes ourselves and register them as in-memory files: DuckDB-WASM's own HTTP reader trips
// on the CDN redirect in some browsers/webviews ("No files found that match the pattern").
// Sources, in order: a same-origin copy (deploy-space.sh puts one in ./data/), then the
// dataset repo. Both end up on the Xet CDN, which intermittently serves a cached
// Access-Control-Allow-Origin for a different origin (Firefox: "does not match"); two
// independent CDN objects give the loader a second chance.
const DATASET = "https://huggingface.co/datasets/jimfhahn/lc-bibframe-hubs/resolve/main/";
const SOURCES = [new URL("./data/", location.href).href, DATASET];
const db = await DuckDBClient.of();
const t0 = performance.now();
const status = display(html`<p>Loading the graph (~97 MB, once)…`);
const isParquet = (b) => b.length > 8 && b[0] === 0x50 && b[1] === 0x41 && b[2] === 0x52 && b[3] === 0x31; // "PAR1"
async function fetchParquet(name) {
  const errors = [];
  for (const base of SOURCES) {
    try {
      const res = await fetch(base + name);
      if (!res.ok) throw new Error(`HTTP ${res.status}`);
      const bytes = new Uint8Array(await res.arrayBuffer());
      if (!isParquet(bytes)) throw new Error("not a Parquet file"); // e.g. dev-server HTML fallback
      return {bytes, base};
    } catch (e) {
      errors.push(`${base}${name}: ${e.message}`);
    }
  }
  throw new Error(`Could not load ${name}:\n${errors.join("\n")}`);
}
async function attach(name) {
  const {bytes, base} = await fetchParquet(name);
  const mb = (bytes.length / 1e6).toFixed(0); // buffer is transferred to the worker (detached) below
  await db._db.registerFileBuffer(name, bytes);
  status.append(html` <small>${name} ${mb} MB ✓${base === DATASET ? " (dataset)" : ""}</small>`);
}
await Promise.all([attach("relations.parquet"), attach("hubs_connected.parquet")]);
await db.query(`CREATE TABLE rels AS SELECT source_id, target_id, rel_type FROM read_parquet('relations.parquet')`);
await db.query(`CREATE TABLE hubs AS SELECT hub_id, title, marc_key, media, agents FROM read_parquet('hubs_connected.parquet')`);
await db.query(`CREATE TABLE deg AS SELECT id, count(*) AS n FROM (SELECT source_id AS id FROM rels UNION ALL SELECT target_id FROM rels) GROUP BY id`);
const counts = await db.queryRow(`SELECT (SELECT count(*) FROM hubs) AS hubs, (SELECT count(*) FROM rels) AS rels`);
display(html`<p>Ready: ${counts.hubs.toLocaleString()} connected Hubs, ${counts.rels.toLocaleString()} relations (${((performance.now() - t0) / 1000).toFixed(1)}s).`);

1. Find a Hub

Type a title or author. Results with the most connections come first, so the original work is usually at the top. Try picking a translation (e.g. Pride and prejudice. Turkish) to see the climb in step 2.

// Deep link: ?hub=<uuid> opens straight on that Hub (used by the VuFind plugin sidebar).
const hubParam = new URLSearchParams(location.search).get("hub");
const q = view(Inputs.text({label: "Search", placeholder: "title or author…", value: hubParam ? "" : "Pride and prejudice", submit: true}));
const chosen = view(Inputs.select(candidates, {
  label: "Hub",
  value: candidates[0],
  format: (d) => `${d.title}  ·  ${d.marc_key ?? "—"}  ·  ${d.degree} links`
}));

2. Climb to the original work

If the Hub you picked is a translation or arrangement, follow its translationof / arrangementof links upward until there is nothing above. That Hub is the family root (the Work, in IFLA-LRM terms) and everything below is drawn around it.

3. The family

Everything one step away from the root, in four groups:

  • Derivative works: adaptations, sequels, parodies, films, operas. The interesting part; open by default and colored by how surprising the relationship is (red = most surprising).
  • Expressions: translations and arrangements. Predictable, so collapsed by default (opened automatically if that is where you started).
  • Series & parts: whole/part, series, continuations, revisions.
  • Other related: links with no specific type.

Click a title to make it the new root. Families overlap, and following a film adaptation into its family is where the unexpected connections show up. Shift-click opens the Hub on id.loc.gov. Choose the indented tree (hierarchy plus columns) or a tidy tree (hierarchy only).

const shown = view(Inputs.checkbox(["Derivative works", "Expressions", "Series & parts", "Other related"], {
  label: "Show", value: ["Derivative works", "Series & parts", "Other related"]
}));
const layout = view(Inputs.radio(["Indented tree", "Tidy tree"], {label: "Layout", value: "Indented tree"}));

4. The same family as a table

Sortable version of the tree (click a column header). The ↗ links go to id.loc.gov.

Notes for extending this notebook

  • db is a DuckDB-WASM client with tables rels(source_id, target_id, rel_type), hubs(hub_id, title, marc_key, media[], agents[]) and deg(id, n). Query it from any cell with db.query(sql, params) or db.sql`…`.
  • family is the list of the root’s one-hop neighbors, each with tier, group, dir, sameAuthor, url; climb holds the root and the path up to it; setFocus(hub_id) re-roots everything.
  • The surprise tiers (cell after the loader) are copied from the VuFind plugin’s RelationshipInferrer.php. Change them there first.
  • Ideas: color by degree, add a second hop for tier-1/2 members, show rel_type_freq.parquet as a rarity bar, or compare two families side by side.

Design notes

Why climb first, then fan out? Arastoopoor (2022, Library Hi Tech 40:1) found that people with no particular book in mind picture a family top-down (work → translations and versions), but people looking for a specific edition or translation navigate bottom-up, from that item to the work and across to its siblings. The page supports both: step 2 is the climb, step 3 the fan-out, and “you started here” marks your entry point.

Why show the relationship type? Pauman Budanović & Žumer (2021, CCQ 59:7) built and tested a cataloging interface on IFLA-LRM entities and the relationships between them, rather than on flat records. Here the relationship (sequel, parody of, film adaptation) gets its own column for the same reason: it is the most informative thing about a family member.

Why an indented tree? Merčun, Žumer & Aalberg’s FrbrVis studies (J. Doc. 2016; JASIST 2017) tested four ways of drawing work families: indented tree, radial tree, circle-pack, and sunburst. The indented tree and the sunburst did best on both speed and user preference. The indented tree also reads like a table, so extra columns fit naturally; the tidy tree is there for comparison.

Surprise tiers. Colors follow the companion VuFind plugin: tier 1 creative transformations (parody, inspiration), tier 2 cross-medium adaptations (film, opera), tier 3 continuations and other adaptations, tier 4 serial/structural links, tier 5 predictable ones (translations, series membership).

Data. jimfhahn/lc-bibframe-hubs, the Library of Congress BIBFRAME Hubs bulk dump (2026-05-05) converted to Parquet; the dataset card documents the schema. Source code for this page and the plugin: jimfhahn/vufind-bf-hubs-plugin.