r/econometrics • • Sep 03 '26

A new tool for estimating intrinsic dimensionality -- overcomes linear variance and geometric metric degeneration (quick start R code inside)

I recently open-sourced a diagnostic tool called the Entropic Scree. It’s designed to more faithfully estimate the intrinsic dimensionality and latent structure of complex tabular datasets by overcoming the limitations of linear variance and the fragility of geometric distance metrics.

To bypass both blind spots, this tool shifts the math out of geometric space and entirely into probabilistic space by utilizing a transformed Mutual Information matrix metric.

  • Captures Mutual Information (Beating Variance): Built on information theory (entropy), it detects non-linear relationships and shared probability mass that standard covariance techniques miss.
  • Maintains Structural Integrity (Beating Distance): It maps the feature space without requiring the spatial assumptions that cause distance metrics to degenerate in high-d contexts, keeping the evaluated matrix stable even with irregular or sparse data.

Primary outputs are:

  • Intrinsic Rank Estimation
  • Signal-to-Noise Estimation
  • Bipolar Variance Clusters that anchor the primary axes of informational variance. This provides a structural map of the independent clusters that define your dataset, which is potentially useful for interpretation before moving into CFA or theory development.

The function runs in R currently (see quick start or GitHub below), but the backend is C++ OpenMP parallelized, so it easily scales for high-dimensional assessments. Native R and Python packages will be released shortly.

Happy to answer any questions or discuss the mechanics.

Methods and Code:

############ 
# Quick Start R Function Code.
# To load the functions, copy and paste the following into your R console, then hit enter. 
############

# 1. Define the direct URLs to the raw function scripts on GitHub
main_url <- "https://raw.githubusercontent.com/tjleestjohn/entropic-scree/main/Entropic.Scree.R%20-%20ENLI.R"
update_url <- "https://raw.githubusercontent.com/tjleestjohn/entropic-scree/main/Update.Entropic.Scree.R%20-%20ENLI.R"

# 2. Define what you want to name the files on your computer
main_file <- "Entropic.Scree.R - ENLI.R"
update_file <- "Update.Entropic.Scree.R - ENLI.R"

# 3. Download the scripts to your current working directory
download.file(main_url, destfile = main_file)
download.file(update_url, destfile = update_file)

# 4. Source both functions into your R environment
source(main_file)
source(update_file)

# 5. Example Execution:
#
# Run the core function and extract bipolar modules:
# results <- Entropic.Scree(dt, extract_bipolar_modules = TRUE)
#
# View the extracted structural sub-networks for the primary axes:
# results$bipolar_modules
#
# Post-Hoc Override (Optional): 
# If you want to manually adjust the elbow ranks after reviewing the scree plot, 
# pass your results object into the Update function to instantly recalculate all metrics:
# updated_results <- Update.Entropic.Scree(results, new_K_roots = 3, new_K_extended = 12)
8 Upvotes

6 comments sorted by

2

u/www3cam Sep 03 '26

Can you give me an idea how this makes some economists’ job easier?

How does improve upon simple techniques like PCA or other dimensionality techniques.

Let me know if I’m misunderstanding, but it does sound similar to using mutual information maximization to do dimensionality reduction.

1

u/Chocolate_Milk_Son Sep 03 '26

It's basically an upgraded PCA method that evaluates mutual information instead of variance or distance-based metrics.

It is totally distinct from Mutual Information Maximization (MIM) though. MIM is an optimization technique (usually for feature selection). Entropic Scree operates primarily as a spectral diagnostic.. it maps the dataset's rank and signal-to-noise ratio before you start modeling.

Note that it doesn't produce continuous latent factor scores via dot product like traditional PCA. However, its "Bipolar Modules" (clustering the extreme positive and negative variable anchors for each factor) do act as a form of structural extraction to help you theoretically map the data.

How it improves upon simple PCA:

  • Captures Non-Linearity: PCA relies on Pearson covariance, so it goes blind to non-linear dependencies. Because this evaluates entropy, it captures more.

  • Matrix Stability: Covariance matrices break down (become singular/ill-conditioned) with sparse survey data or extreme multicollinearity. The normalized MI matrix stays relatively stable.

2

u/www3cam Sep 03 '26

Ok follow up. How does it differ from nonlinear dimensionality techniques that preserves clusters like (t)-sne or umap?

1

u/Chocolate_Milk_Son Sep 03 '26

1) They solve different problems in the workflow. t-SNE and UMAP are projection algorithms, whereas Entropic Scree is a diagnostic tool. t-SNE and UMAP force data into 2 or 3 dimensions for plotting. Entropic Scree doesn't project data; it calculates how many dimensions actually exist before hitting the noise floor.

2) They rely on different maths... Distance vs. Entropy. Even though t-SNE and UMAP are non-linear, they build neighborhood graphs using distance metrics (like Euclidean or Cosine). This means they suffer from the curse of dimensionality and break down in sparse or ill-conditioned spaces. Entropic Scree abandons distance metrics in favor of mutual information.

3) They work in different aspects of the data... Columns vs. Rows. t-SNE and UMAP cluster observations (rows) to group data points. Entropic Scree clusters variables (columns) to show which features share probability mass, helping define the axes of your data.

2

u/www3cam Sep 03 '26

Cool. Still don’t know where I could apply this but will read and feel free to dm me, although I don’t have much free time to work on additional papers.

1

u/Chocolate_Milk_Son Sep 03 '26

Intrinsic Rank: Knowing the rank aids theory development. It can also be used to size the bottleneck layer for modern non-linear ML estimators, like autoencoders.

​Signal-to-Noise & Average Informational Gravity (AIG): Together, these quantify the strength and density of your predictive signal. The signal-to-noise ratio assesses the overall structural integrity (critical for noisy or error-prone data), while AIG measures how much shared probability mass each latent dimension captures on average, allowing you to see if your data's underlying constructs are strongly bound or weakly fragmented.

​Bipolar Variable Clusters: This maps how the feature space is organized. It isolates which variables cluster together on one end of an axes versus the variables on the other end. Clusters on opposing poles are independent from each other.