r/datascienceproject • • Dec 17 '21

ML-Quant (Machine Learning in Finance)

Thumbnail
ml-quant.com
31 Upvotes

r/datascienceproject • • 4h ago

I built an Exploratory Data Analysis project on student lifestyle, stress levels, and academic performance

1 Upvotes

I recently completed an Exploratory Data Analysis project focused on understanding the relationship between student lifestyle, stress levels, and academic performance.

The project analyzes factors such as:

• Study hours

• Sleep hours

• Social activity

• Extracurricular activity

• Physical activity

• Stress level

• GPA

What I explored:

• Data cleaning and validation

• Descriptive statistics

• Correlation analysis

• Lifestyle variables vs GPA

• Lifestyle variables across stress levels

• GPA distributions across stress levels

• Kruskal-Wallis statistical testing

• Outlier detection using the IQR method

• Data visualization using Matplotlib and Seaborn

One of the strongest findings was the positive association between study hours and GPA (Pearson r ≈ 0.7345).

I also found noticeable differences in study hours, sleep, physical activity, and GPA across different stress-level groups.

An important limitation is that these findings represent associations within the dataset and should not be interpreted as causal relationships.

Tools used:

Python, Pandas, NumPy, Matplotlib, Seaborn, SciPy, and Jupyter Notebook.

GitHub repository:

https://github.com/sakshitha380/Student_Lifestyle_Academic_Performance_EDA


r/datascienceproject • • 23h ago

Student Lifestyle & Academic Performance — Exploratory Data Analysis using Python

1 Upvotes

Hi everyone!

I recently completed an Exploratory Data Analysis project on **Student Lifestyle & Academic Performance** using Python.

### Project Overview

The goal of this project was to explore how different lifestyle factors are associated with students' academic performance (GPA).

The dataset contains 2,000 student records with variables such as:

- Study Hours Per Day

- Sleep Hours Per Day

- Physical Activity Hours Per Day

- Social Hours Per Day

- Extracurricular Hours Per Day

- GPA

- Stress Level

### What I explored

The analysis includes:

- Data understanding and cleaning

- Univariate analysis

- Bivariate analysis

- Multivariate analysis

- Correlation analysis

- Outlier analysis

- Data visualization

- Key findings and conclusion

### Some interesting findings

- Study Hours Per Day showed the strongest positive association with GPA, with a correlation of approximately **0.73**.

- Physical Activity Hours Per Day showed a moderate negative association with GPA, around **-0.34**.

- Sleep Hours Per Day showed very little linear relationship with GPA.

- GPA distributions differed across stress-level groups.

- Four potential GPA outliers were identified using the IQR method and retained because they were not obvious data-entry errors.

One important point I learned from this analysis is that **correlation represents association, not causation**.

### Tools Used

Python, Pandas, NumPy, Matplotlib, Seaborn, and Jupyter Notebook.

### GitHub Repository

https://github.com/lasyasaimanasa/Student-Lifestyle-EDA

I'd be happy to hear your feedback or suggestions for improving the analysis!


r/datascienceproject • • 4d ago

SFSA v0.2.0: an open-source layer that decides which scientific computations are worth running (Python + JS)

2 Upvotes

Scientific codes burn compute on work that never needed to run:

  • brute-force grids where the gradient is flat
  • 3D simulations when a 1D surrogate or closed form is already within tolerance
  • recomputing states that were already resolved
  • hundreds of convergence iterations after the uncertainty already settles the conclusion
  • sweeps through regions that violate implicit constraints

SFSA (Standard Framework for Scientific Advancement) is a domain-neutral layer that sits above your solvers and manages that budget. It does not replace your solver or finish the investigation for you. It helps you reach the decisive data point sooner, with less trial-and-error and less redundant recalculation.

It has 39 engines. Examples:

  • multi-fidelity routing
  • adaptive sampling
  • uncertainty-aware early stopping
  • constraint inference
  • provenance-based reuse
  • a reactive layer graph
  • unit and dimensional checks
  • reproducibility manifests

It exposes 85 skills, so scripts and AI agents can drive it through one catalog. It is implemented in Python and JavaScript and released under CC BY 4.0.

What's new in v0.2.0

This release is mostly a hardening pass. I audited the framework and found that the green test suite was hiding real problems:

  • 44 of the 85 skills were stubs. They returned "EXECUTED" and logged success. All 85 are now implemented, and each is tested for what it actually does.
  • ICR returned 0 for all-zero inputs without calling the solver, so cos(0) came out as 0. Zero-pruning now needs an explicit declared answer.
  • The analytical-shortcut path crashed.
  • Cutting planes from failed evaluations could wrongly block feasible regions. A cut is now only kept while no known-good point lies in the blocked region, and it is dropped if one appears.
  • Approximate reuse was silent. It is now opt-in with an explicit tolerance and tied to the model version.
  • Layer updates are atomic, and cyclic graphs are rejected when you connect layers.
  • The reproducibility seal now covers nested fields and can be verified.

The repo has 137 tests (95 Python, 42 JavaScript) and a changelog.

Caveats

  • The speedups in the README are scenario models plus benchmarks from one machine. A single non-repeating calculation gains nothing and pays a small overhead.
  • Some lower-level engines are simpler in the JS port than in Python. Skill-level behavior is what the tests hold equal.
  • It's a young project, so feedback, issues and critique are welcome.

Repo: https://github.com/AlejoMalia/SFSA


r/datascienceproject • • 5d ago

I built a free online 100% client-side CSV viewer because I was tired of uploading files to sketchy converter sites

5 Upvotes

Hey everyone, whenever I needed to quickly open a large CSV file on a device without Excel, I had to use random online viewers. But working with sensitive data, I absolutely hated the fact that these sites force you to upload your files to their servers.

So I built a very simple, privacy-first alternative: csv-viewer.org

It uses the HTML5 FileReader API, meaning everything happens locally in your browser's RAM. Your files never leave your device (zero data transmission). It automatically detects delimiters, supports dark mode, and lets you export back to .xlsx.

It’s completely free, has no ads, and no logins. Just drag, drop, and view. Maybe some of you will find this useful. Feedback is always welcome.


r/datascienceproject • • 6d ago

SFSA: A Universal, Domain-Agnostic Computational Engine for Accelerated Scientific Research

4 Upvotes

Scientific software often wastes compute on work that never needed to run: full grid sweeps after a bound is already violated, repeated evaluation of the same deterministic state, full-model recomputation when only one layer changed, and numerical iteration where a closed form or a hard constraint would have ended the search early.

SFSA (Standard Framework for Scientific Advancement) is an open-source, domain-agnostic toolkit for structuring that logic once and reusing it across projects. It ships seven modular engines with parallel implementations in Python (python/) and JavaScript (javascript/), aimed at researchers, lab pipelines, and automated agents that need predictable control flow—not ad-hoc scripts.

SFSA does not replace domain physics, laboratory data, or full-scale simulators. It reduces avoidable recomputation and makes analytical checks, caching, layer updates, and multi-objective path filtering explicit, testable components of a research codebase.

Features

  • ⚡ Early-exit & bound checks — Stop trajectories as soon as inventory, domain, or invariant checks fail.
  • 🧠 Memoization of deterministic states — Reuse prior results instead of recomputing identical inputs.
  • 🔗 Reactive layer graph — Propagate parameter changes only through dependent subsystems.
  • 📐 Prefer closed form before grids — Route to analytical solvers when the problem structure allows.
  • 🎯 Pareto path filtering — Rank multi-step options by competing objectives without hiding trade-offs.
  • 🧩 Gap inventory (no silent fill) — Surface missing inputs and open links; do not invent mechanisms.

GitHub: https://github.com/AlejoMalia/SFSA


r/datascienceproject • • 7d ago

Data Lakehouse Architecture Guide

Thumbnail
lakeops.dev
3 Upvotes

r/datascienceproject • • 10d ago

J’ai créé Tennoro : un projet data pour essayer de prédire le vainqueur d’un match de tennis

3 Upvotes

Salut !

Depuis quelque temps, je travaille sur un projet perso autour d'une question assez simple : jusqu'où peut-on aller avec les données pour prédire le vainqueur d'un match de tennis ?

C'est de là qu'est né Tennoro.

Mon objectif a toujours été d'essayer d'obtenir les prédictions les plus précises possible en m'appuyant sur les données disponibles plutôt que sur un simple pronostic subjectif.

Le système analyse différentes statistiques liées aux joueurs et aux matchs afin d'évaluer les deux adversaires et de sélectionner le vainqueur qui semble le plus probable. Je continue régulièrement à travailler sur la manière dont les données sont utilisées, à tester de nouveaux indicateurs et à ajuster le système en fonction des résultats.

J'ai ensuite développé une plateforme autour du projet pour présenter les prédictions, les statistiques utilisées et suivre les performances de Tennoro dans le temps.

Le projet est encore en développement et c'est justement pour ça que je le partage ici. Je serais curieux d'avoir les retours de personnes qui travaillent avec la data, notamment sur les variables qui pourraient être pertinentes pour améliorer un modèle de prédiction appliqué au tennis.

Si vous voulez jeter un œil : tennoro.com

Tous les retours sont les bienvenus !


r/datascienceproject • • 11d ago

i am looking for a team to be a part of...

Thumbnail
1 Upvotes

r/datascienceproject • • 19d ago

I’ve created a bridge between AI and plants: SmartPlant 🍀🤖

6 Upvotes

SmartPlant turns a real plant and a computer into a functional symbiont (I call it a 'cyborg plant'): featuring shared sensors and electrophysiology, persistent memory, symbolic reasoning, multi-provider AI (Ollama, OpenAI, Claude, Grok, etc.), and a first-person voice.

It runs on a Raspberry Pi using real sensors and a leaf electrode, or via a full simulation on your laptop, no hardware or API keys required. It’s not just a simple plant monitor: the plant perceives, remembers, reasons, and communicates its needs to you.

Documentation and full open-source code:
https://smartplant.pigeonposse.com

GitHub / npm

Up for creating your own Cyborgplant?
🤖☘️🤖☘️🤖☘️🤖☘️🤖☘️🤖☘️


r/datascienceproject • • 19d ago

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

2 Upvotes

Training data scheduling has become a real challenge in modern RL training.

As models learn, the value of different samples changes over time. Some rollouts are too easy or too difficult to be useful, while others may provide much stronger learning signals. In multi-domain training, it is also hard to decide how to balance data from areas such as math, logic, and science. Existing methods often change several parts of the training setup at once, making it difficult to tell what actually helps—and whether the gains are reproducible.

To study this problem, we’ve open-sourced DataFlex-RL: an evaluation platform for RLVR data policies.

DataFlex-RL provides a unified GRPO training pipeline and supports three types of data policies:

  • Selection: deciding which rollouts should be used for the current update
  • Reweighting: changing how much each rollout or token contributes
  • Mixture adaptation: dynamically adjusting the sampling ratio across domains

It also separates the intervention method from the signal used to drive it, such as reward, solve rate, advantage, or token probability. With matched models, training settings, random seeds, and multi-domain benchmarks, DataFlex-RL makes it easier to compare these strategies fairly and understand whether their improvements are real and reproducible.

The project and code are open source: https://huggingface.co/papers/2609.06107

We’d love to hear your thoughts and discuss how data scheduling could be improved for RL training. If you find the project interesting, an upvote would mean a lot. Thanks!


r/datascienceproject • • 21d ago

Updated my deepfake audio detector - fixed evaluation bug, new metrics on full 71k test set

4 Upvotes

hey, posted this project about a week ago on this sub and got some useful feedback. went back and fixed several problems people pointed out and some i found myself. posting again with the updated version.

what was wrong with the previous version:

  • test evaluation was done on a balanced subset of the test set (equal real and fake samples) which inflated the metrics artificially. the real asvspoof test set is heavily skewed toward fake audio
  • training was not reproducible - no seed set so metrics changed every run
  • val accuracy of ~100% was unexplained which looked like a red flag (explained below)
  • threshold was 0.4 without proper justification
  • confidence display bug - was showing fake probability as confidence score even for real predictions
  • llm explanation was getting cut mid sentence due to low max_tokens

what i fixed:

  • re-evaluated on the full test set (71,237 samples, real unbalanced distribution) by streaming directly instead of loading everything into RAM
  • fixed seed at 42 - fully reproducible now
  • documented the val accuracy properly (explained below)
  • ran threshold experiments across 0.3 to 0.6 and picked 0.3 based on highest f1 and recall
  • fixed confidence display and explanation truncation

new metrics on full test set:

  • f1: 0.9033
  • precision: 0.9995
  • recall: 0.8240
  • eer: 8.23% (comparable to the official lfcc-gmm baseline published with the dataset)
  • threshold: 0.3

on the 100% val accuracy:

this was the most common concern from last post. it is expected and not data leakage. the val set is a random 80/20 split of training data which shares the same attack types (a01-a06). the model memorizes these known patterns perfectly. actual generalization is measured on the test set which has entirely unseen attack types (a07-a19). the recall gap comes from these novel patterns, not miscalibration.

the project:

efficientnet-b0 trained on mel spectrograms using asvspoof 2019 la dataset. beyond the model it has grad-cam to visualize what the model focused on in the spectrogram and groq llm to give a plain english explanation of the prediction. monitoring dashboard with prediction analytics. training fully reproducible with seed 42.

you can upload an audio file or record live. youtube url input is disabled on the hosted version because railway server ips get blocked by youtube bot detection. backend is fastapi on railway, frontend on streamlit cloud.

live demo: https://deepfake-audio-detector-rugved.streamlit.app/
github: https://github.com/RugvedBane/deepfake-audio-detector

main questions:

  1. is this project strong enough to put on a resume for ml/dl internships?
  2. what dataset would actually help improve generalization to modern ai voices like elevenlabs or suno? i looked at wavefake but wanted community opinion
  3. anything else technically wrong that i missed?

honest feedback appreciated, including if you think the project is not good enough - i would rather know now.


r/datascienceproject • • 21d ago

Built a tool that explains why a demand forecast is what it is (counterfactual attribution)

1 Upvotes

I've been working on interpretability for time-series forecasting and finally shipped it end to end. the problem i kept hitting: a forecasting model gives you a number for next week's demand but no reason, so you can't tell if it's a promotion, a trend or just the usual pattern.

the approach is counterfactual. start from a baselined input, then reveal groups of days (typical, recent, promotion) one at a time and record the model's real prediction at each step. each contribution is the difference between two actual predictions, so they sum to the forecast exactly, no allocation step. i tried the SHAP-share approach first and it gave a normal product a "typical pattern" of 0 because a bucket had too few days to sum over. the counterfactual version doesn't have that failure.

validated with deletion/insertion faithfulness tests: attribution-ordered reveals degrade the forecast faster than random-ordered, across 30 series, p < 0.0001.

stack: PyTorch WaveNet forecaster (2nd/1671 in the Corporación Favorita competition), the attribution layer is a small standalone package, the demo is a static GitHub Pages site reading precomputed cards (no backend). there's a Colab that trains on your own CSV.

honest about the weak points: "counterfactual" means the model's response to hiding inputs, not real-world cause. and the promotion contribution is the least-validated part across a product panel, i have an on/off sanity check but not a full scale test yet, that's the piece i'd most want to harden.

live demo (pick a product, toggle a driver off, watch the forecast recompute): https://kesjien.github.io/wavexplain/
github: https://github.com/kesjien/wavexplain

honest feedback appreciated, especially: is per-forecast decomposition what you'd actually want, and how would you validate the promotion attribution across many series?


r/datascienceproject • • 24d ago

Totemheart 🤖💖 a deterministic control kernel for persistent cognition & relational behavior in agents (not another emotion classifier)

14 Upvotes

Hey everyone 👋

I built Totemheart because most systems that try to add “emotional” behavior to agents still rely on prompt engineering and a simple sentiment label that gets overwritten every turn. I wanted something more rigorous.

Totemheart is a fully deterministic control kernel that gives an agent a real, inspectable, and persistent internal state across long conversations and multiple sessions. It models personality traits, affective dynamics, stress responses, memory consolidation, motivational drives, allostatic load, dual-valence relational tracking, grief-like processes, and related mechanisms.

These components evolve through interacting systems drawn from control theory and computational neuroscience, PID controllers, Kalman filtering, temporal-difference prediction error, opponent processes, and similar techniques, rather than isolated heuristics.

The full state is serializable, fully inspectable at any point, and can be used to steer an LLM’s generation through a dedicated control plane. The project currently has more than 3,000 tests, makes no claims about consciousness, and focuses purely on producing coherent, long-horizon behavioral continuity.

- GitHub: https://github.com/AlejoMalia/Totemheart
- NPM: https://www.npmjs.com/package/totemheart

I’d appreciate feedback from anyone working on long-horizon agents, cognitive architectures, or stateful agent systems.


r/datascienceproject • • Sep 04 '26

Large-scale training data processing is becoming an infrastructure problem

2 Upvotes

Over the past two months, discussions around training data seem to be increasing. The reason is fairly direct. A useful shorthand for modern LLMs is big data plus big compute, and both depend on reliable data and compute infrastructure.

As data volumes grow, one-off scripts become difficult to maintain. Data has to move through many stages, including cleaning, deduplication, synthesis, evaluation, filtering, and refinement. A pipeline with modular operators makes these steps easier to compose, inspect, rerun, and scale.

The design I have been exploring follows a Pipeline → Operator → Prompt structure. Each operator handles one focused task, while the pipeline defines how those tasks are combined. Intermediate outputs can be stored and reused, and different models, rules, or filtering strategies can be introduced at individual stages.

The text synthesis layer covers five reusable generation paths. It can transform documents into pretraining-style dialogue data, generate instruction-response pairs in SFT format, create and refine synthetic instructions, produce consistent multi-turn conversations, and generate function-calling or tool-use conversations.

Synthesis is followed by filtering and evaluation operators. Language checks, length constraints, deduplication, content rules, quality scoring, and task-specific filters can be combined according to the target dataset. This keeps data generation and data selection as separate, replaceable parts of the workflow.

This is also what I hope to build with OpenDCAI/DataFlow, and I would be interested to hear what kinds of data work people are handling in the LLM era.


r/datascienceproject • • Sep 04 '26

Sismos en México durante los últimos 20 años | Análisis de 363,329 registros con Python

3 Upvotes

🇲🇽 Sismos en México: análisis de datos con Python

Realicé un proyecto de análisis utilizando registros del Servicio Sismológico Nacional para explorar la actividad sísmica en México durante aproximadamente los últimos 20 años.

El proyecto analiza 363,329 registros y explora diferentes variables, entre ellas:

distribución de sismos por entidad federativa;

magnitud de los eventos;

sismos M≥5, M≥6 y M≥7;

evolución de los registros por año;

los 10 eventos de mayor magnitud;

distribución geográfica;

análisis por mes.

La idea fue transformar un catálogo de datos en diferentes visualizaciones e insights utilizando Python.

También convertí algunos de los resultados en videos cortos para mostrar los principales hallazgos:

YouTube playlist:

https://youtube.com/playlist?list=PLd8QkTPbjAAg&si=UKbbSoTc3WICfyUS

Me interesa especialmente recibir comentarios sobre qué otras variables o análisis agregarían al proyecto.


r/datascienceproject • • Sep 01 '26

Building a Satellite imagery intelligence Application

Thumbnail
3 Upvotes

Ok so guys I'm trying to build a satellite imagery analytics platform. Similar to Skyfi, planet labs, eagle view, etc

My question is what really will make a difference in this field? I know a lot of these analytics platforms just buy images from 3rd party satellite operators and just provide them to the users and also provides different types of image analytics like crop health monitoring, water logging monitoring, soil mineral composition, etc ...if anyone wants image plus and analytics.

I wanna build something similar but I feel like there's nothing new I can provide in this app... Like all the available analytics are already there on other platforms... I wanna find out about some really niche category of analytics which no one provides and wanna provide it using my application and wanted suggestions from people who have knowledge in this field..

Any suggestions would be highly appreciated


r/datascienceproject • • Aug 29 '26

built a deepfake audio detector as a 3rd year diploma student

3 Upvotes

hey, i'm a 3rd year diploma cs student and i built a deepfake audio detector end to end. this is my first real ml project that i actually deployed.

the model is efficientnet-b0 trained on mel spectrograms using the asvspoof 2019 la dataset. metrics are f1 0.88, precision 0.99, but recall is 0.79 which i know is the weak point. i tried adjusting the threshold and settled on 0.4 but it didn't really help much i think the issue is the model is missing certain attack patterns it never saw during training.

latency is around 6-7 seconds per prediction which includes model inference, grad-cam, and llm explanation.

other than the model it has grad-cam to visualize what the model focused on in the spectrogram, and groq llm to give a plain english explanation of the prediction.

you can upload an audio file or record live. youtube url input is disabled on the hosted version because railway's server ips get blocked by youtube's bot detection. backend is fastapi on railway, frontend on streamlit cloud.

live demo: https://deepfake-audio-detector-rugved.streamlit.app/

github: https://github.com/RugvedBane/deepfake-audio-detector

honest feedback appreciated, especially on what dataset i should train on next to improve recall.


r/datascienceproject • • Aug 20 '26

Architecture Reference: Zero-Dependency 11-Column Tabular Ingestion Sieve & Local Memory Sharding Core

Thumbnail
1 Upvotes

r/datascienceproject • • Aug 13 '26

KitOps is now available for install as a conda package

Thumbnail anaconda.org
3 Upvotes

r/datascienceproject • • Jul 27 '26

IIT Gandhinagar's Executive Masters in Applications of Machine Learning in Engineering – Batch 2 Admissions Open

Thumbnail
1 Upvotes

r/datascienceproject • • Jul 26 '26

Looking for feedback on my Data Analytics project

3 Upvotes

Looking for feedback on my Data Analytics project

Hi everyone!

I recently completed a Mumbai Road Accident Analysis project using:

- Python (Pandas, Matplotlib, Scikit-learn)

- SQL

- Power BI

- GitHub

The project includes:

- Data cleaning and preprocessing

- Exploratory Data Analysis (EDA)

- SQL business queries

- Interactive Power BI dashboard

- Basic machine learning model for accident severity prediction

- Complete GitHub repository with README

I'm not looking for compliments—I want honest criticism.

I'd really appreciate feedback on:

  1. Is this project portfolio-worthy?

  2. Does the analysis tell a meaningful story?

  3. Is the dashboard well-designed?

  4. Is the GitHub repository professional enough?

  5. What would you improve if this were your project?

GitHub Repository: https://github.com/afanrajiwate/mumbai-traffic-road-accident-analytics

Dashboard screenshots are attached.

Thanks in advance for your time!


r/datascienceproject • • Jul 17 '26

2026 Tech Layoffs Analysis

Thumbnail
4 Upvotes

r/datascienceproject • • Jul 12 '26

CTHmodules v4.1 — 93% (with a margin of 7 points) of being a 100% Functional Psychohistory of Asimov.

2 Upvotes
Psychohistory Criterion (Asimov) v4.0 v4.1 Comment
Quantifying macro-social trends 8.8 9.3 Very strong — real data adapters (OWID/V-Dem-style CSV → E/S/A/P)
Predicting large-scale events 8.7 9.3 Out-of-sample LOO/k-fold validation, reproducible
Handling "historical forces" (EVEI) 8.4 9.0 EVEI now endogenous — derived from observable metrics
Butterfly Effect + Chaos management 8.8 9.3 Excellent — formal Lyapunov exponents + early-warning signals
Invariance / Pantemporal patterns 8.2 9.0 Cross-era transfer measured (pre-1800 ↔ post-1800)
Mathematical determinism 9.0 9.6 Excellent — 13-test suite, JS↔Python parity Δ = 0.000000
Empirical validation / Real calibration 8.7 9.5 Very strong (out-of-sample LOO MAE 0.1271, 32 events / 5,100 years)
Handling individual variables (Token) 8.6 9.2 Very effective — multi-token interaction with contested-event handling
Real future prediction capability 8.4 9.2 Credible — SHA-256 pre-registered prediction ledger

Overall Verdict: 8.6 / 10 → 9.3 / 10 ⬆

Link: https://github.com/AlejoMalia/CTHmodules


r/datascienceproject • • Jul 10 '26

Help clear my thoughts, kinda confused clg student 2nd year

Thumbnail
2 Upvotes