network_data.js · 6.25 MB · generated 2026-10-06

A citation map of research on modularity in biology

207 hand-picked “highly relevant” papers, each scored on five ways of thinking about modules, plus the thousands of papers they cite or are cited by.

207seed papers
14,315papers in total
21,276links between them
5scoring axes
Each dot is one seed paper. Hover or tap a dot for its title and scores.

Explore the interactive visualization

This report is accompanied by an interactive visualization of the same data. Open it alongside the report to explore the papers on your own.

Open the interactive visualizationOpens in a new tab

Where it started: the scoring spreadsheet

Before there was a visualization, there was highly_relevant_axis_scores.csv: one row per paper, with the five scores and the evidence behind them. network_data.js is built from it, and everything about the five axes comes from this sheet.

209rows
21columns
207distinct papers
77.6 KBfile size

The 21 columns, in four groups

Identity

paper titleDOI

The only link to OpenAlex. The DOI is how the script later finds each paper’s year, venue and citations.

Final scores

trait_scorenetwork_scorefunction_scoreevolution_trajectory_scorehierarchy_score

Integers from 0 to 100, no empty cells. These are the values the map uses.

Keyword evidence

traits_keyword_scorenetworks_keyword_scorefunctions_keyword_scoreevolution_keyword_scorehierarchy_keyword_score…_terms (5 columns)

The automatic first pass: a score per axis and the search patterns that matched. Many term cells are empty because nothing matched (172 of 209 for traits).

Provenance and notes

notesauto_primaryauto_hypertext_source

A one-sentence justification per row (50–208 characters), the automatic main axis, and which text was read: abstract (103 rows), PDF (75) or bibliography (31).

How much the reading changed the keyword scores

Each paper has two numbers per axis: what the keyword scan produced and the final score. They differ in 439 of the 1,045 cells (42%). Only 13 papers kept every keyword score untouched. Scores went up 363 times and down 76 times; 59 keyword hits were reduced to zero and 70 zeros were raised after reading.

Share of papers whose score changed

Out of 209 rows. Network and function were corrected most often.

Average score, before and after

keyword scanfinal score

The keyword scan and the final score agree most closely on trait and hierarchy (correlation 0.87) and least on evolution (0.75).

The automatic main axis (auto_primary) only ever takes three values: networks (107 rows), functions (84) and traits (18). Evolution and hierarchy are never picked automatically. It matches the highest final score in 165 of 209 rows, so the manual adjustment changed which axis leads for about one paper in five. The auto_hyper column marks only 7 papers, all as a network and function blend.

What the JavaScript file adds

The spreadsheet has no years, venues, authors, citation counts or links between papers. Those come from OpenAlex through the DOI. The script also adds the standardized pentagon coordinates, and it brings in about 14,000 cited and citing papers as context. The sections below describe that result.

What network_data.js is

The file packages a literature review as data. At its center are 207 papers chosen as “highly relevant” to one question: how are biological systems organized into semi-independent parts, or modules? Each of those papers carries a score on five meanings of the word module (trait, network, function, evolution & trajectory and hierarchy), plus a short note explaining the score.

The papers range from brain networks and gene regulation to skull shape, synthetic gene circuits and even modular deep learning. The question is the same in each: what counts as a module, and how do modules arise and work? Around these papers the file adds their surroundings: the roughly 14,000 works they cite or that cite them, and the links between all of them. That lets you see which papers share foundations and where the five meanings of modularity meet.

The four top-level keys

KeyContents
metaCounts, the axis names, and the constants used to place papers on the map.
nodes14,315 papers. Each is a compact record with an OpenAlex ID, title, year, venue, first author, citation count and so on.
edges21,276 links, stored as [source, target, type] using positions in the node list.
seeds207 richer records: the five axis scores, matched keywords, reviewer notes and where the scoring text came from.

Seed papers versus context papers

The 207 seeds (flag s = 1) carry the full scoring fields. The other 14,108 nodes are “context” papers with basic metadata only. Seven of those were added manually and are all from 2026.

One seed, Simon’s The Architecture of Complexity (1962), carries a note explaining that OpenAlex only has later reprints, so its year and venue come from the original publication.

How the 207 seeds were scored

Every seed receives a score from 0 to 100 on five axes. Each axis is a different meaning of the word “module”. Together they let you ask which kind of modularity a paper is really about.

1. Keyword scanTerms such as “community structure” or “composable” are counted in the text and turned into a first score per axis.
2. ReadingThe text comes from the PDF (74 papers), the abstract (102) or the bibliography (31).
3. AdjustmentA written note explains each final score. 26 notes flag a keyword hit as spurious; 13 say the keywords were right.
4. MappingScores are standardized and projected onto a pentagon, giving each paper an x/y position.

The map

The five axes point to the corners of a pentagon, clockwise from the top: trait, network, function, evolution & trajectory, hierarchy. A paper’s position is the sum of its above-average scores along those directions, so papers drift toward the corner they emphasize and sit between corners when they blend ideas.

The stored coordinates (pc) match this construction exactly, and pw records the paper’s strongest standardized score. The metadata also stores a value of 0.685 labelled pca2d, which appears to be the share of variance the 2-D layout keeps.

Which axis dominates

Papers by their highest-scoring axis. Network and function together account for about 80%.

Papers that blend two ideas

Pairs of axes both scoring 50 or more. Most papers (131) are strong on one axis only; 65 are strong on two or more.

What the corpus looks like

Seed papers by year

Two older seeds (1962 and 1999) are left out of the chart. Output peaks in 2015–2019.

Where they are published

VenueSeeds
Nature Communications13
PLoS Computational Biology11
Science8
PNAS7
bioRxiv7
Scientific Reports6
Molecular Systems Biology6

113 distinct venues overall. 182 articles, 15 reviews and 10 preprints. OpenAlex assigns the seeds to 66 different topics, led by bioinformatics and genomic networks (44) and gene regulatory network analysis (40).

The most cited seeds

PaperYearCitations
The Structure and Function of Complex Networks200318,901
Community structure in social and biological networks200215,802
Modularity and community structure in networks200612,525
Network Motifs: Simple Building Blocks of Complex Networks20027,580
The Architecture of Complexity19624,487

The median seed has 96 citations, so a handful of classics dominate the totals.

The links between papers

Each edge has a type code in its third position. There are three kinds.

18,224

Citation (type 0)

A seed cites a reference (13,531 links to 10,136 papers) or a later paper cites a seed (4,693 links from 3,972 papers). Context papers are never linked to each other.

476

Seed to seed (type 1)

One seed cites another. 197 seeds take part. The most cited from inside the set is From molecular to modular cell biology (58 times).

2,576

Bibliographic coupling (type 2)

Two seeds share references. A fourth value stores how many, from 2 to 38 (mean 4.0). The strongest pairs share 38 references.

The references that many seeds lean on show the field’s common ground: Cytoscape, Network biology: understanding the cell’s functional organization and the Louvain community-detection paper are each cited by 20 or 21 seeds.

Of the 14,108 context papers, 12,003 touch only one seed, so the outer ring is mostly sparse. Most are recent: about two thirds date from 2010 onward.