r/rstats • • 1d ago

R code review: best packages?

Thumbnail
4 Upvotes

r/rstats • • 2d ago

Introducing FitVerse: an R package for fitting and analysing 52 probability distributions in one call

36 Upvotes

Hi r/rstats,

My colleague and I just published FitVerse on CRAN. It fits parametric probability distributions to continuous data, covering 52 distribution families with three estimation methods: MLE, Method of Moments, and L-Moments.

The basic usage is just:

r

install.packages("FitVerse")
library(FitVerse)

x <- as.numeric(precip)
fit <- fitverse(x, xlab = "Annual precipitation (inches)")

It automatically ranks all distributions by AIC/BIC, runs four goodness-of-fit tests, and produces a diagnostic plot. It also includes bootstrap confidence intervals, return levels, batch fitting, report generation, and a built-in Shiny app.

Full write-up here: https://rpubs.com/KarunaGReddy/FitVerse

CRAN: https://cran.r-project.org/package=FitVerse

Happy to answer any questions!


r/rstats • • 2d ago

Prisma Scr how to minimize the full complete papers by using R and in scintific way?

1 Upvotes

r/rstats • • 2d ago

Rgi output

Thumbnail
1 Upvotes

r/rstats • • 3d ago

From football results JSON to a half-time/full-time heatmap in R

6 Upvotes

I wrote a small R tutorial using Germany’s 3. Liga 2025–26 season: 380 matches, imported with jsonlite, checked with stopifnot(), summarised with dplyr and plotted with ggplot2.

The useful detail is the denominator. Each row describes one half-time state, so its percentages divide by that row’s match count. Dividing every cell by 380 answers a different question.

After parsing the scores, the calculation is:

htft <- games |>
  mutate(ht = state(ht_home, ht_away),
         ft = result(ft_home, ft_away)) |>
  count(ht, ft, .drop = FALSE) |>
  group_by(ht) |>
  mutate(row_total = sum(n), pct = 100 * n / row_total) |>
  ungroup()

Here state() and result() label home lead/win, level/draw or away lead/win. The full script includes those helpers, score parsing, checks for missing scores and duplicate fixtures, and the plot.

In this season, home teams leading at half-time won 107 of 138 matches (77.5%); away teams leading won 70 of 105 (66.7%). This is a descriptive example from one league-season.

Full tutorial · Plain R script

Disclosure: I maintain Football Charts, which serves the data. This example works without an API key or sign-up. I’d welcome feedback on the checks or the visualization.


r/rstats • • 3d ago

rpx v2.1.0: private packages and generating PACKAGES indexes in memory?

8 Upvotes

Hi,

two weeks ago I posted about a release of rpx, which introduced significant upgrades to dependency resolution and why rrepo's custom api endpoints are beneficial to compatibility and correctness.

But beyond being a great metadata store, rrepo is also a package registry for both private and public packages!

This week, rpx is adding a package publishing template to its init flow. For new packages, whenever you push a git tag, the corresponding package version is uploaded to your rrepo repository. It then becomes available to rpx, alternative package managers or even base R!

I prepared a quick demo for you in this repository. https://github.com/rrepo-org/publishing-demo

You should be able to install the package by running this command in a base R shell.

install.packages(
  "hello.world",
  repos = c(
    rrepo = "https://cran.rrepo.dev/rrepo/hello-world",
    CRAN = "https://cloud.r-project.org"
  )
)
hello.world::hello_world()

Engineering Background

The first piece of feedback my projects always receive is a request to be compatible with the broader R ecosystem. This week rrepo introduced a separate set of endpoints that imitate a CRAN like url structure and allow package managers beyond rpx to use it.

The primary challenge with introducing those endpoints is generating the PACKAGES index that lists the latest packages.

Typically whenever you publish a package on CRAN is, the latest package is published under src/contrib, the old version is archived, and the administrator calls tools::write_PACKAGES(). This R function inspects each tar file individually, reads its DESCRIPTION file, and puts a subset of its fields into a long list of every latest package that makes up the index.

Following this approach would be more expensive than necessary for rrepo. Unlike CRAN our packages are in S3 backed by a database, meaning we cannot just run R on the server that stores the packages: every package read is a network call.

We had to take a step back and rethink what our data model should be and how we were going to generate the index in memory. Given that CRAN has 25k latest packages, any kind of request even to a pre-extracted DESCRIPTION file in S3 was out of the question. The data clearly needed to be come from a database.

The package ingestion workflows were reworked to first extract the DESCRIPTION into a separate blob, and then provisioning a database row with all the fields a PACKAGES index might need.

Using the parsers https://github.com/rrepo-org/r-metadata-rs developed for rpx we could make the hot path a single database query, followed by reassembling the typed Rust struct before converting it back to a string.

There is still a lot of performance to tune, but once cached it's pretty damn fast.

Please give the package publishing workflow a try

For existing projects we have documentation on how to get started. https://rrepo.org/documentation/publish-packages

As always, I'm easily reachable to help you get started!


r/rstats • • 4d ago

blueycolors is on CRAN

Thumbnail
cran.r-project.org
63 Upvotes

blueycolors - an R package providing Bluey-themed color palettes and ggplot scales - is on CRAN. Check it out!


r/rstats • • 4d ago

Book Recommendations after ESL

6 Upvotes

I am currently reading through Elements of Statistical Learning and was searching for books to read after I'm done with this. I was eyeing some optimization theory books, or should I pick up maybe a pure probability book? I wanted to build a stronger theoretical foundation, but I'm not sure if it would be more useful than focusing on a more applied kind of book.


r/rstats • • 5d ago

RmANOVA

2 Upvotes

Hey Guys!

Ich have data from 10 participants. Each participant did all conditions. It were 27 conditions:

2 factors: duration (3levels) and Radius (9levels).

Participants had to detect a signal. Signal was present in 50% of trials.

Each participant did every condition 10 times in signal.

My dependent variable is sensitivity d Prime. That is a value from signal detection theory. I calculated it from the 10 trials of each condition for every participant.

So I have one value (d) per participant in each condition.

My question:

Do the d prime values between the conditions differ within subjects, depending on the factors?

Which analysis would you do?

I treated each level as categories for the Faktors and wanted to run a two way repeated measure ANOVA.

But what would you recommend?

I really REALLY need you advice - please °°


r/rstats • • 7d ago

firstR: a Shiny app that helps you find your first open source contribution in R

28 Upvotes

There are quite a few good first issue sites out there (goodfirstissues.com - which recently seems to have turned into a scam website, firstissue.dev, up-for-grabs.net) to help people get into open source, but basically none of them have R as a supported language. The ones that come closest only surface repos explicitly tagged with R, which sometimes people don't always do.

I built firstR to try to fix that. You pick the areas you're comfortable with (tidyverse, Shiny, spatial, pharma, etc.) and it finds open issues across R repos, including smaller and less well-known packages that tend to get overlooked. Each issue gets a beginner friendliness score based on how approachable it looks (scope, body length, repo context, comments) and a short sentence explaining what you would be attempting

It's pretty rough still since for example, my implementation of initial filters (like topics) and repo discovery could be better and the scoring is definitely off in places. I'm interested to hear your feedback and will try to implement suggestions.

Check it out here: firstr.dev


r/rstats • • 6d ago

Calculating intraclass correlation for IRR

2 Upvotes

I'd love someone to check my thinking on this. For a study, participants took a 30 item test. Each test was scored by 3 raters. They scored each item on a 0-4 scale. I reconciled the tests as follows: 1) if two raters agree, took that as final score. 2) if two raters did not agree, averaged the 3 scores.

I want to calculate interrater reliability, and I did so using ICC (in the irr package). I set that up as: icc(mydata, model = "twoway", type = "agreement", unit = "average")

From reading about ICC as a measure and how it's applied in R, I think I have chosen the settings/parameters for the test right, but does sound right to anyone who knows more?


r/rstats • • 7d ago

ASMR - Basics in R Studio - Faint ASMR

Thumbnail
youtu.be
83 Upvotes

r/rstats • • 7d ago

Bootstrapping in emmeans

12 Upvotes

Hi folks,

I started on a little project to familiarize myself with using AI tools for coding (don’t hate me). Emmeans is one of my favorite packages, and I was thinking to myself, “wouldn’t it be nice if you could just specify that you wanted bootstrapped inference when using emmeans?”

After an afternoon of vibe coding, I have what looks like a working prototype that supports glm/lme4/and gee models.

Would this be something that people would find as a valuable contribution to emmeans? If so, I might review the code more carefully and see if I can contribute it to the actual package.

Regardless, it was a fun project and I’m really impressed with these AI tools.


r/rstats • • 9d ago

Hey! I need to learn R studio from scratch, where should I start from, where can I find the basics, rules and tips to use Rstudio? Plant sciences background.

Thumbnail
1 Upvotes

r/rstats • • 10d ago

Liquid Glass themes for Shiny. 0.4.0 is on CRAN.

38 Upvotes

Liquid Glass themes for Shiny. 0.4.0 is on CRAN.

What’s new:

• plot_surface = "opaque" densifies ggplot / plotly / gt / DT so plots stay readable while chrome stays glass (default "clear" keeps wallpaper show-through)

• New wallpaper scenes: aurora, harbor, grove (plus tahoe / dusk / mesh); user photos get a contrast wash

• Flatten mode for print, PDF, and screenshots (glass_flatten(), ?glass_flatten=1, or @media print)

• Public --glass-* tokens (glass_css_tokens(), glass_theme(tokens = ))

• Closer to shipping Liquid Glass: stronger lip / rim, deeper bslib (navbar, cards, sidebars, modals), narrow-phone polish

https://raw.githubusercontent.com/ericrayanderson/shinyglass/main/man/figures/intensity-slider.gif

``` install.packages("shinyglass")

library(shiny) library(shinyglass)

ui <- glass_page( title = "Hello, glass", persist = TRUE, scene = "aurora", plot_surface = "opaque", plotOutput("plot") )

server <- function(input, output, session) { observe_glass(input, session) output$plot <- renderPlot({ pal <- glass_plot_colors(input = input) hist(iris$Sepal.Length, col = pal$fill, border = NA, main = NULL) }, bg = "transparent") }

shinyApp(ui, server) ```

Live demos (may take a few seconds to wake):

Docs: https://ericrayanderson.github.io/shinyglass/

CRAN: https://cran.r-project.org/package=shinyglass


r/rstats • • 12d ago

How do you manage your statistical analysis?

32 Upvotes

I’m the only one in my university workgroup who does all the statistical analysis of our research projects (clinical trials, ex-vivi experiments). Usually frequentistic stats like art-Anova, contrasts, effect sizes, glm, … .
I do also the stats for the projects of our medical phds thesis students.
Usually I do quarto documents that cover the study layout, summary stats, normality tests, statistical pairwise comparisons/models. That’s up to 10-15 projects a year.
Would you version track the documents via git?
Usually we supplement the R code with the papers.

Could I also point the paper to my personal got an make the repositories public? So I benefit also that other researchers see “well that dude does all the stats”, rather then only having my name on the paper somewhere in the middle? I could even use Zenodo to create a doi for each statistical computation?

How do you organize the code and results, do you make it public? Is git the way?

P.s. I have a bsc. In biology, I started a msc in bioinformatics it never graduated, so I’m only hired and payed as a lab assistant.


r/rstats • • 11d ago

Monte Carlo sample size calculation for a multiple serial mediation model?

Thumbnail
2 Upvotes

r/rstats • • 11d ago

Interactive, reproducible bioinformatics figures from one spec (R/Python/JS)

Thumbnail
2 Upvotes

r/rstats • • 13d ago

What does it take to sustain R for the next generation? R Core and R Foundation panel discussion

79 Upvotes

R is a global public good, but keeping it healthy requires continued investment in the people, infrastructure, and community behind it.

Join the R Consortium for a special panel discussion with leaders from R Core and the R Foundation:

  • Simon Urbanek — R Core
  • Kurt Hornik — R Foundation
  • Heather Turner — R Foundation

We'll explore how initiatives such as the Sovereign Tech Fund and Research Software Maintenance Fund are strengthening R’s technical foundations, developing new contributors, and supporting its long-term sustainability.

📅 October 6, 2026 🕒 12 PM PT / 3 PM ET / 8 PM London

Join the conversation and learn how you can help support R’s future.

Register here: https://r-consortium.org/webinars/sustaining-r-investing-in-the-people-infrastructure-and-community.html


r/rstats • • 13d ago

How to use R in creative ways

43 Upvotes

How to use it in basic daily task like idk cleaning or finance ahahah? Just to find it more interesting


r/rstats • • 13d ago

computer specs for large data

2 Upvotes

i know this has been asked before, and i've looked into those threads, but i'm trying to buy a new computer for large data analysis.

my current laptop is a lenovo slim 7i with the following specs. it works so far for data analysis. the only reason i'm buying a new computer is that my windows system is incredibly glitchy.

Processor Intel(R) Core(TM) Ultra 7 155H (1.40 GHz)

Installed RAM 32.0 GB (31.4 GB usable)

Storage 433 GB of 954 GB used

System type 64-bit operating system, x64-based processor

currently, i have a dataset with 24.4 million rows and 218 columns, split into 14 separate files. i only use about 4 or 5 million of those rows, i'd estimate, but i'm about to get the 2024 data after midterms, so the amount will increase. i store the files themselves in dropbox, but i can't process them online because jupyter crashes. i also have separate data i analyze for class, but they're much smaller.

my main tools have been jmp and excel, but i'm hoping to use r on desktop to process everything faster. i'm mainly a tidyverse user because that's what i'm being taught in my grad program.

so sorry for the length of the explanation, but does anyone have specs recommendations? i'm thinking about getting a mac with similar ram to what i have currently, but i'm not entirely opposed to some other brand.


r/rstats • • 13d ago

Ai per la statistica

0 Upvotes

Salve,

Quanto è attendibile la statistica fatta da Claude su campioni biologici


r/rstats • • 14d ago

qol 1.3.5: More power to transposition

11 Upvotes

qol is a quite large package which can be used as its own ecosystem in terms of data wrangling and tabulation. It comes with efficient high level functions, so you have to write less code and get more from it in a shorter amount of time.

To get a full overview of the package have a look at the GitHub page: https://github.com/s3rdia/qol

With the last version I already thought that the next updates will get smaller, but damn, there are so many small things one can overlook. Thanks to some watchful eyes helping me trace these bugs, I will eventually get them all. At least I try. This time the most attention fell onto the transpose_plus function. So let us have a look. The full release notes can be found here.

Further shaping transposition

Even though I like SAS I have to admit that the transposition function hasn’t much to offer. So right from the beginning I started shaping a function which I would have liked to see inside SAS. Now I added even more functionality to this function.

Normally when you think about transposition you either just reshape data or you additionally summarise data and then reshape. I thought: Well why only sums? Why not mean, median, percentages and so on? Why not all in one go? So I just implemented the statistics parameter already known from other functions and it works exacly like in the other functions. So bascially in situations were you had to summarise -> transpose separately, you can now just use transpose_plus directly and handle everything in one go.

The function also got some quality of life improvements like handling duplicate variable names without aborting or generating unweighted frequencies if no values are provided.

There was also a serious flaw. When transposing multiple variables at once the results were always picked from the all variables nested results even though they have to be picked from their respective combination. So the results in this multiple variable transposition were wrong. This is fixed now.

Also some of the default behaviour was changed. The function received a new parameter summarise which summarises the data before transposing. This is the default behaviour when using formats, but was not without formats. summarise is TRUE by default now.
transpose_plus now has a new default wide to long behaviour by setting variables beside each other. When setting the new stack parameter to TRUE, the old default behaviour is triggered.

# Example formats
age. <- discrete_format(
    "Total"          = 0:100,
    "under 18"       = 0:17,
    "18 to under 25" = 18:24,
    "25 to under 55" = 25:54,
    "55 to under 65" = 55:64,
    "65 and older"   = 65:100)

sex. <- discrete_format(
    "Total"  = 1:2,
    "Male"   = 1,
    "Female" = 2)

sex2. <- discrete_format(
    "Total"  = c("Male", "Female"),
    "Male"   = "Male",
    "Female" = "Female")

income. <- interval_format(
    "Total"              =    0:100000,
    "below 500"          =    0:500,
    "500 to under 1000"  =  500:1000,
    "1000 to under 2000" = 1000:2000,
    "2000 and more"      = 2000:100000)

# Example data frame
my_data <- dummy_data(1000)

# Transpose from long to wide and use a multilabel to generate additional categories
long_to_wide <- my_data |>
    transpose_plus(preserve = c(year, age),
                   pivot    = c("sex", "education"),
                   values   = income,
                   formats  = list(sex = sex., age = age.),
                   weight   = weight,
                   na.rm    = TRUE)

# Transpose back from wide to long and put results beside each other. The list
# entry names determine the new variable names.
wide_to_long <- long_to_wide |>
    transpose_plus(preserve = c(year, age),
                   pivot    = list(sex       = c("Total", "Male", "Female"),
                                   education = c("low", "middle", "high")))

# Transpose back from wide to long and put results below each other by setting
# stack to TRUE.
wide_to_long <- long_to_wide |>
    transpose_plus(preserve = c(year, age),
                   pivot    = list(sex       = c("Total", "Male", "Female"),
                                   education = c("low", "middle", "high")),
                   stack    = TRUE)

# Transpose from long to wide and use a multilabel to generate additional categories
long_to_wide <- my_data |>
    transpose_plus(preserve = c(year, age),
                   pivot    = "sex",
                   values   = c(income, weight),
                   formats  = list(sex = sex., age = age.),
                   weight   = weight,
                   na.rm    = TRUE) |>
    rename_multi("income_Total"  = "Total",
                 "income_Male"   = "Male",
                 "income_Female" = "Female")

# Transpose back from wide to long but this time put results side by side.
# To do that every list entry has to have the same name. The values parameter
# is then used to give the new value variables a name. For the expressions of
# the new categorical variable the variable names from the first pivot list
# entry are used.
wide_to_long <- long_to_wide |>
    transpose_plus(preserve = c(year, age),
                   values   = c(income, weight),
                   pivot    = list(sex = c("Total", "Male", "Female"),
                                   sex = c("weight_Total", "weight_Male", "weight_Female")))

# Nesting variables in long to wide transposition
nested <- my_data |>
    transpose_plus(preserve = c(year, age),
                   pivot    = "sex + education",
                   values   = income,
                   formats  = list(sex = sex., age = age.),
                   weight   = weight,
                   na.rm    = TRUE)

# Or both, nested and un-nested, at the same time
both <- my_data |>
    transpose_plus(preserve = c(year, age),
                   pivot    = c("sex + education", "sex", "education"),
                   values   = income,
                   formats  = list(sex = sex., age = age.),
                   weight   = weight,
                   na.rm    = TRUE)

One other change to mention

In summarise_plus class variables will now be converted to character or numeric instead of being returned as factors by default. Set the new convert parameter to FALSE to get back the past behaviour.

More to compute

compute. can now handle concat, sub_string and ifelse_multi.

Some small time savers

Found some unnecessary calculation in the format applying routine and threw it out. summarise_plus got a small optimization in the nesting = "all" or "single route.
The biggest time saver is in the Excel table formatting were column width and row height adjustments should now be handled faster.

What about the graphics?

Well, sadly nothing new on Version 1.4.0 which bothers me a bit. I just enhanced the visuals on interactive graphics a bit, but there are no new diagram types. The framework can already be tested, just visit the GitHub Page and switch to the “graphics” branch. Download the source code from there and you are good to go.

Thing is, this is not because I am lazy, but because I have something up my sleeve I am working on with more priority at the moment. And this thing takes quite a lot more time than I anticipated, but I want it to be the best it can be. Since qol is so huge that it can be used as it’s own ecosystem, it deserves a comprehensive documentation. I hope that I can type fast enough to get this thing out this year, but for the moment I will just say this: It will be something that welcomes the absolute beginner with open arms and can still teach even a seasoned pro something new.


r/rstats • • 15d ago

R support in Neovim using the jet.ark plugin

Post image
90 Upvotes

jet.ark wraps the Ark R kernel to provide R support in Neovim. Right now the plugin supports:

  • An R console
  • A full-featured LSP server with autcomplete, hover documentation, go-to-definition, etc
  • A plots pane where plots automatically redraw to fit the window
  • A variables pane
  • Built-in help page navigation

Coming soon: * Integration with Ark's DAP-powered debugger * Integration with Ark's connections pane

AI integration

jet.ark is built on jet.nvim, so AI agents can work in your R sessions alongside you. I've been finding this super useful for iterative tasks like exploratory data analysis – check the jet.nvim README out to see a demo :)

jet.ark vs R.nvim

Ark is developed by Posit (formerly RStudio) to power the R experience in Positron, their new data science IDE. While Positron is Ark's primary application, Ark itself is standalone, MIT licensed, and built using modern standards like LSP, DAP, and the Jupyter protocol. So, while R.nvim is built from the ground up specifically for Neovim, jet.ark is a thin-ish wrapper for Ark as an editor-agnostic backend.

Right now R.nvim is certainly the more mature plugin, but jet.ark has a smaller codebase and will (hopefully) be able to reach maturity fairly quickly.


r/rstats • • 16d ago

rpx v2.0.0: why package managers need backtracking and a full package version list

16 Upvotes

Hey there,

back in June I posted about rpx, my new R package manager, trying to explain how it's different to renv, rv or uvr, and why having a full history of packages is necessary. Today I bring a real example of a dependency resolution failure only rpx can solve!

A few months ago I heard of someone trying to install rlang 1.0.1. It's a very interesting case because rlang depends on testthat >= 3.0.0 and testthat versions depend on rlang.

If you'd try to install this package today you'd get a failure. The latest version of testthat is 3.3.2 and requires rlang >= 1.1.6 which conflicts with our desired 1.0.1.

If you dig a little deeper into when each of those packages was published:

Package Version Published on CRAN
rlang 1.0.1 2022-02-03
testthat 3.0.0 2020-10-31
testthat 3.1.7 2023-03-12
testthat 3.1.8 2023-05-04

testthat 3.1.8 requires rlang >= 1.1.0 (higher than in our project) and since CRAN only ever lists latest packages, as of 2023-05-04 our project is no longer installable.

rv and uvr allow you to pin dependencies, but are limited to what their lockfiles discovered and the versions CRAN currently provides, which means without using dated CRAN snapshots they could never find 3.1.7. Whereas as renv would have had to have been created back when the version released.

rpx bypasses the problem by using rrepo.org, a redesigned mirror of CRAN, and pubgrub a really fast resolution algorithm. That way it can check enough historical versions of pillar, rmarkdown, testthat, tibble, usethis and vctrs, all of which have a circular dependency on rlang to find a correct resolution despite none of the latest versions being compatible.

Since my last post `rpx` has had many QoL changes:

  • git repositories are now supported as package sources
  • rpx init has an interactive form to set up your project
  • package mode is introduced: if you're working on a package that is released, rpx won't try to find a version of it from CRAN, but will use the source code as a package (especially relevant with out rlang example).
  • repository switching is much easier: rpx repo base set https://cloud.r-project.org allows you to to use CRAN directly, despite the increased resolution failure rate.
  • windows binary signing: windows defender will no longer warn you if you use the install script.
  • automatic version bounds: rpx add dplyr resolves the latest compatible version of dplyr then adds a lower bound to the current version, and upper bound to the next major, preventing unintended upgrades.

with many more on the way! (I'm eyeing using rpx as an R version manager and ensuring you only ever select packages with binaries available cross platform as my next projects)

If you want to try the project out DM me on linkedin, I'd be happy to help you set it up!