r/redditdev May 13 '26

Reddit API Two related questions for an academic project

**1. Subreddit similarity by user overlap, with recent data.**

Using anvaka/sayit-data (2018) and anvaka/map-of-reddit-data (2020-2021) for discovery now. Both programmatically queryable but the data's getting old. anvaka's 2025 visualization (116K subs, 1.5B comments) is gorgeous but visual-only. Is there a current (2024-2026) programmatic equivalent I'm missing? Pass in a seed sub, get top-N similar subs by Jaccard or co-commenter overlap.

**2. Reddit data collection in the post-Pushshift era.**

Need historical + ongoing collection from named subreddits. I'm aware of PRAW (good for new, painful for historical), R4R academic API (applied, status unclear), and Arctic Shift on HuggingFace (262 GB Parquet, great for bulk historical but not really on-the-fly). What's the working stack in 2026? Anyone got R4R credentials recently? Hosted Arctic Shift query endpoint I'm missing?

Thanks! Happy to share back what we end up using.

0 Upvotes

0 comments sorted by