← Saúl Huitzil / Projects & reports / Publications
207 hand-picked “highly relevant” papers, each scored on five ways of thinking about modules, plus the thousands of papers they cite or are cited by.
This report is accompanied by an interactive visualization of the same data. Open it alongside the report to explore the papers on your own.
Open the interactive visualizationOpens in a new tabBefore there was a visualization, there was highly_relevant_axis_scores.csv: one row per paper, with the five scores and the evidence behind them. network_data.js is built from it, and everything about the five axes comes from this sheet.
The only link to OpenAlex. The DOI is how the script later finds each paper’s year, venue and citations.
Integers from 0 to 100, no empty cells. These are the values the map uses.
The automatic first pass: a score per axis and the search patterns that matched. Many term cells are empty because nothing matched (172 of 209 for traits).
A one-sentence justification per row (50–208 characters), the automatic main axis, and which text was read: abstract (103 rows), PDF (75) or bibliography (31).
Each paper has two numbers per axis: what the keyword scan produced and the final score. They differ in 439 of the 1,045 cells (42%). Only 13 papers kept every keyword score untouched. Scores went up 363 times and down 76 times; 59 keyword hits were reduced to zero and 70 zeros were raised after reading.
Out of 209 rows. Network and function were corrected most often.
The keyword scan and the final score agree most closely on trait and hierarchy (correlation 0.87) and least on evolution (0.75).
The automatic main axis (auto_primary) only ever takes three values: networks (107 rows), functions (84) and traits (18). Evolution and hierarchy are never picked automatically. It matches the highest final score in 165 of 209 rows, so the manual adjustment changed which axis leads for about one paper in five. The auto_hyper column marks only 7 papers, all as a network and function blend.
The spreadsheet has no years, venues, authors, citation counts or links between papers. Those come from OpenAlex through the DOI. The script also adds the standardized pentagon coordinates, and it brings in about 14,000 cited and citing papers as context. The sections below describe that result.
The file packages a literature review as data. At its center are 207 papers chosen as “highly relevant” to one question: how are biological systems organized into semi-independent parts, or modules? Each of those papers carries a score on five meanings of the word module (trait, network, function, evolution & trajectory and hierarchy), plus a short note explaining the score.
The papers range from brain networks and gene regulation to skull shape, synthetic gene circuits and even modular deep learning. The question is the same in each: what counts as a module, and how do modules arise and work? Around these papers the file adds their surroundings: the roughly 14,000 works they cite or that cite them, and the links between all of them. That lets you see which papers share foundations and where the five meanings of modularity meet.
| Key | Contents |
|---|---|
meta | Counts, the axis names, and the constants used to place papers on the map. |
nodes | 14,315 papers. Each is a compact record with an OpenAlex ID, title, year, venue, first author, citation count and so on. |
edges | 21,276 links, stored as [source, target, type] using positions in the node list. |
seeds | 207 richer records: the five axis scores, matched keywords, reviewer notes and where the scoring text came from. |
The 207 seeds (flag s = 1) carry the full scoring fields. The other 14,108 nodes are “context” papers with basic metadata only. Seven of those were added manually and are all from 2026.
One seed, Simon’s The Architecture of Complexity (1962), carries a note explaining that OpenAlex only has later reprints, so its year and venue come from the original publication.
Every seed receives a score from 0 to 100 on five axes. Each axis is a different meaning of the word “module”. Together they let you ask which kind of modularity a paper is really about.
The five axes point to the corners of a pentagon, clockwise from the top: trait, network, function, evolution & trajectory, hierarchy. A paper’s position is the sum of its above-average scores along those directions, so papers drift toward the corner they emphasize and sit between corners when they blend ideas.
The stored coordinates (pc) match this construction exactly, and pw records the paper’s strongest standardized score. The metadata also stores a value of 0.685 labelled pca2d, which appears to be the share of variance the 2-D layout keeps.
Papers by their highest-scoring axis. Network and function together account for about 80%.
Pairs of axes both scoring 50 or more. Most papers (131) are strong on one axis only; 65 are strong on two or more.
Two older seeds (1962 and 1999) are left out of the chart. Output peaks in 2015–2019.
| Venue | Seeds |
|---|---|
| Nature Communications | 13 |
| PLoS Computational Biology | 11 |
| Science | 8 |
| PNAS | 7 |
| bioRxiv | 7 |
| Scientific Reports | 6 |
| Molecular Systems Biology | 6 |
113 distinct venues overall. 182 articles, 15 reviews and 10 preprints. OpenAlex assigns the seeds to 66 different topics, led by bioinformatics and genomic networks (44) and gene regulatory network analysis (40).
| Paper | Year | Citations |
|---|---|---|
| The Structure and Function of Complex Networks | 2003 | 18,901 |
| Community structure in social and biological networks | 2002 | 15,802 |
| Modularity and community structure in networks | 2006 | 12,525 |
| Network Motifs: Simple Building Blocks of Complex Networks | 2002 | 7,580 |
| The Architecture of Complexity | 1962 | 4,487 |
The median seed has 96 citations, so a handful of classics dominate the totals.
Each edge has a type code in its third position. There are three kinds.
A seed cites a reference (13,531 links to 10,136 papers) or a later paper cites a seed (4,693 links from 3,972 papers). Context papers are never linked to each other.
One seed cites another. 197 seeds take part. The most cited from inside the set is From molecular to modular cell biology (58 times).
Two seeds share references. A fourth value stores how many, from 2 to 38 (mean 4.0). The strongest pairs share 38 references.
The references that many seeds lean on show the field’s common ground: Cytoscape, Network biology: understanding the cell’s functional organization and the Louvain community-detection paper are each cited by 20 or 21 seeds.
Of the 14,108 context papers, 12,003 touch only one seed, so the outer ring is mostly sparse. Most are recent: about two thirds date from 2010 onward.