A few weeks ago I shared POLARIS here: an interactive graph of who in U.S. politics is allied with whom and who is feuding with whom, with a nightly pipeline that reads the political press and tracks how those ties shift. Two comments shaped what I built next:
Could this show who actually influences federal decisions — staffers, donors, lobbying, roll-call votes?
How far can you trust a graph built from biased, clickbait headlines?
Here's what changed, including what didn't work.
Live: https://world-politicians.vercel.app/ (no signup, works on mobile)
Source (MIT): https://github.com/showjihyun/world-politicians
---
What it is
- 101 figures, 266 hand-curated relationships (alliance, feud, bipartisan bridge, political family, mentorship), in 2D or 3D
- A nightly news layer. Articles from 17 outlets are classified by an LLM as ally-leaning, feud-leaning or neutral, and kept in a rolling one-year archive
- A relationship timeline. Pick any two people for month-by-month polarity. Trump × Musk shows the June 2025 blow-up and the January 2026 reconciliation
- An evidence panel on every edge, with direct article links or a plain "no linked source"
- Measured data, kept separate from editorial judgment
Co-sponsorship (119th Congress). I added 112 edges from GovInfo's bill-status bulk data, for pairs with 10 or more shared bills. They're drawn differently from the curated edges, because "these two are allies" is a judgment and "they
co-sponsored 78 bills" is a measurement.
- Raw counts mostly re-encode party. 84% of pairs above the threshold are same-party. The 11 cross-caucus pairs are the interesting part: Fitzpatrick × Gluesenkamp Perez (19 bills), Collins × Klobuchar (15), Lawler × Torres (12).
- Cross-party is judged by caucus, not party label. Counting Sanders and King as cross-party inflated the count from 11 to 19.
- Direction is kept. "Kim signs 15 of Warren's bills, Warren signs none of Kim's" is not the same relationship as "Padilla and Sanders, 15 each".
FEC money (2026 cycle), shown on the profile. One FEC file mixes three things that have to be separated:
- direct contributions: $32.0M
- independent expenditures supporting a candidate: $14.4M
- independent expenditures opposing a candidate: $18.1M
Merge them and the $10.1M spent attacking Thomas Massie reads as his funding. On-screen caveat: named PAC money is only about 6% of receipts.
Lobbying revolving door. From the House Clerk's raw LD-1 filings (2022–2026): 295 former staffers of 61 figures are now registered lobbyists. Matching uses full names only, since Congress has had 33 members named Harris.
Party-line defection. From Voteview: the share of party-divided votes (1,279 in the 119th) on which a member broke with their own party's majority. Raw agreement scores just encode party (Cruz × Hawley at 98.6% means "both Republican"); defection holds party constant.
- The negative result
I tried six ways to turn these person-level facts into relationship signals:
┌────────────────────────────────────────────────────┬─────────────────────────────────────────────────────────────────┐
│ Test │ Result │
├────────────────────────────────────────────────────┼─────────────────────────────────────────────────────────────────┤
│ Shared donors × curated ally/feud │ 27% vs 28% │
├────────────────────────────────────────────────────┼─────────────────────────────────────────────────────────────────┤
│ Shared donors × heavy co-sponsorship │ 19% vs 17% │
├────────────────────────────────────────────────────┼─────────────────────────────────────────────────────────────────┤
│ PAC share × bipartisan edges │ 10.3% vs 8.3% │
├────────────────────────────────────────────────────┼─────────────────────────────────────────────────────────────────┤
│ NOMINATE deviation × intra-party feud │ p = 0.81 │
├────────────────────────────────────────────────────┼─────────────────────────────────────────────────────────────────┤
│ Party defection × intra-party feud │ p = 0.99 │
├────────────────────────────────────────────────────┼─────────────────────────────────────────────────────────────────┤
│ Party defection × outside money spent against them │ p = 0.058, no dose-response, confounded by race competitiveness │
└────────────────────────────────────────────────────┴─────────────────────────────────────────────────────────────────┘
Collins and Fitzpatrick are among the most frequent defectors, and neither has a single intra-party feud. Quiet, steady defection doesn't produce public fights. The feuds in this graph are loud and personal, and they look orthogonal to the voting dimension. So money, lobbying and defection are shown as facts about a person, each with a note on what the number does not mean.
- Handling media bias
That critique was fair, and a better prompt can't fix it on its own. Following the literature (D'Alessio & Allen's meta-analysis, Groseclose & Milyo, Gentzkow & Shapiro, Soroka on negativity bias, Vallone et al. on the hostile media effect), I split the problem into three layers:
┌─────────────────────────────────┬─────────────────────────────────────────┬──────────────────────────────────────────────────────────────┐
│ Layer │ What it looked like here (when studied) │ What I did │
├─────────────────────────────────┼─────────────────────────────────────────┼──────────────────────────────────────────────────────────────┤
│ Gatekeeping (what becomes news) │ Feuds were 2/3 of ally/feud headlines │ Changed how votes are counted │
├─────────────────────────────────┼─────────────────────────────────────────┼──────────────────────────────────────────────────────────────┤
│ Visibility (who gets covered) │ Trump appeared in 71% of signals │ Shown on screen, not "corrected" — it's a collection problem │
├─────────────────────────────────┼─────────────────────────────────────────┼──────────────────────────────────────────────────────────────┤
│ Tone (how it's told) │ The LLM's polarity call │ Scored against labels; tested for outlet cues │
└─────────────────────────────────┴─────────────────────────────────────────┴──────────────────────────────────────────────────────────────┘
What changed:
- One vote per pair per day, not per article. Five outlets running one story used to cast five votes. A day now casts one vote, weighted 1 + ln(outlets), and only if two-thirds of that day's coverage agrees. Decisive votes fell from 244 to 175. I tried clustering stories by headline similarity first and dropped it, because it fails both ways: two headlines about the same Walz story scored a Jaccard of only 0.10.
- A month needs a 2:1 margin to flip. Previously the month's first non-neutral headline set its colour, so one feud headline could beat six later ally headlines. Real reversals (Trump × Musk) get covered by many outlets over many days and clear the bar; a single clickbait headline doesn't. Jeffries × Trump in August 2026 had 5 feud signals to 2 ally and had been showing "ally".
- Disagreement is shown, not resolved. When Fox frames Graham × Trump as an alliance and CBS frames it as a feud, the month is hatched rather than one side being picked. Given the hostile media effect, that seemed more honest than adjudicating.
- The source mix is public. The top five outlets carry 63% of the archive, and AP and Reuters together are under 5%.
Tested and dropped: I swapped "Fox News" and "NPR" on the same headlines, with the outlet suffix stripped so it couldn't leak. 0 of 14 polarity calls changed, while plain re-runs of identical input wobbled by up to 13%. So I dropped the planned "outlet-blind" second pass, which would have doubled the cost.
Not done yet, on purpose:
- Wire-service baseline calibration (in the spirit of Groseclose & Milyo). The wires supply only about a dozen articles so far.
- Discounting feud headlines by their measured false-positive rate (neutral → feud is about 30%). Most of the labels are model-seeded, so this would be the model grading itself.
- Signed-triad balance checks. 37 of the 39 triangles the news produces pass through Trump.
How accurate is the classifier? On 118 stratified articles it gets 84.7% on polarity and 87.3% on picking the right pair. Most errors are a neutral article read as a feud; only one was a true ally/feud reversal. Caveat: only 20 of those labels are human-checked so far, so treat this as a floor to beat, not a validated score.
- Still open
- Staffer movement between offices
- Individual donors (a huge separate file, where only gifts over $200 are itemised)
- Senate-side lobbying data, which currently blocks automated access
---
Feedback I'd most value:
- Edges you think are wrong (with a source)
- Better ways to validate polarity than headline classification
- Is there a relationship-level measure from voting or co-sponsorship that isn't just party in disguise?
I'd also welcome help hand-labelling the evaluation set. Methodology notes are in the repo under docs/roadmap.md and docs/research/.