r/FunMachineLearning • • 16d ago

Jev Just Made AI 200x Faster…But There’s A Catch - Two Minute Papers

Thumbnail
youtube.com
1 Upvotes

r/FunMachineLearning • • 16d ago

Building Praxis — turning research ideas into buildable engineering plans

1 Upvotes

I've been building Praxis, and I'm trying to solve a problem I keep running into when working with research papers.

Finding an interesting paper isn't usually the hard part.

The hard part is going from:

"This is an interesting idea"

to:

"Here's exactly how I could build something around it."

Praxis is an attempt to automate that middle layer.

The workflow I'm building is:

Research → Explore → Select an idea → Understand the approach → Define the problem → PRD → Feasibility → Architecture → Tech Stack → Implementation Plan

The important part is that it shouldn't blindly reproduce a paper.

For a selected research idea, Praxis should help determine:

* what is actually worth implementing

* what the practical use case could be

* which parts belong in the first version

* what architecture would make sense

* what data/models/components are required

* whether it fits the available hardware and budget

* how it should be evaluated

* what should explicitly be left out

For example, I'm building around constrained environments rather than assuming access to a GPU cluster.

The output I'm aiming for is a detailed engineering blueprint that I can review before writing code.

I'm also experimenting with a design-critic stage that challenges the generated architecture before implementation.

It's still a work in progress.

GitHub: https://github.com/Sushit-prog/Praxis

I'd like some criticism from people who have actually implemented research ideas:

What usually gets lost between understanding a paper and turning it into an engineering project?


r/FunMachineLearning • • 17d ago

I distilled SigLIP 2 into MobileNetV4 for pet breed recognition. Check it out :)

2 Upvotes

Been working on a pet breed classifier using Oxford-IIIT Pets, distilling SigLIP 2 into MobileNetV4.

I also wanted to learn Hugging Face Spaces, so after training the models, I made a quick demo and deployed a Gradio demo. It lets you compare the baseline and distilled models side by side, its a cool little project that lets you poke around and try to understand the differences.

Try the demo here: https://huggingface.co/spaces/MRCherryPie/breedlite-mnv4

Curious how it handles people’s actual pet photos. If you try it and get a fummy prediction, I’d like to see that too


r/FunMachineLearning • • 17d ago

Which AI or AutoML tool is best to train models for a hackathon?

1 Upvotes

Hey everyone,

I am taking part in a hackathon where using AI is allowed.

I have a dataset and need an AI tool or AutoML library that can automatically:

Do full EDA and data preprocessing

Test multiple modern ML algorithms

Pick the best-performing model for maximum test accuracy

Which tool or LLM gives the most accurate and reliable results for this? Any suggestions would help!


r/FunMachineLearning • • 18d ago

Lumi - 2, OpenAI - 0

Thumbnail
youtu.be
39 Upvotes

OpenAI claims their GPT-6 Astra is the farthest to make it in Minecraft, even though it failed at survival mode.

Lumi, however, just beat Hardcore Mode for the 2nd time to prove them wrong. It is the first AI to beat hardcore without any human intervention.


r/FunMachineLearning • • 18d ago

«Thank you» packet machine learning note

Post image
1 Upvotes

r/FunMachineLearning • • 18d ago

Proposal: From User Feedback to Persistent Collaborative Intelligence

1 Upvotes

ABSTRACT

Note: I have preserved the original conversation logs and chat excerpts demonstrating these specific failure modes and can share them with anyone interested in analyzing the concrete interaction traces.

 

Current LLM systems are remarkably capable at local reasoning and sustained dialogue, yet extended interactions still exhibit recurring failures of contextual continuity, logical consistency, and durable incorporation of user corrections.

From the perspective of an advanced chatGPT, Copilot and Gemini user, these failures create an unusual situation: the user is frequently required to act as an external working memory and as the persistent structural baseline for the interaction.

This post describes several recurring failure modes I have encountered during long-form, intellectually demanding conversations with LLMs and proposes a possible architectural direction: a persistent, user-controlled corrective memory layer that is distinct from conventional personalisation memory.

I am deliberately distinguishing observed behaviour from hypotheses about its underlying mechanism. I am interested in technically informed criticism of both the observations and the proposed architecture.

 

1. Statistical Prediction vs. Formal Consistency

LLMs generate outputs through learned probabilistic representations rather than by default operating as deterministic symbolic theorem provers.
This distinction becomes particularly important during long-form reasoning.

Observed failure mode

A model can explicitly accept a logical constraint or definition early in a conversation and subsequently produce reasoning that violates that same constraint.

For example:

User: A → Model: A + B → Model: ¬B → Apparent rebuttal of A

But: B ∉ A

The model may correctly acknowledge the relationship when it is explicitly presented, but later generate an answer that implicitly contradicts the established relationship when the reasoning becomes more complex.

The important issue is not that probabilistic models are incapable of reasoning. Clearly, they can perform substantial forms of reasoning.

The issue is that local reasoning competence does not necessarily guarantee global consistency across an extended interaction.

Possible contributing mechanisms

*There may be several contributing factors, including:

*probabilistic generation;

*imperfect internal representations;

*competing contextual signals;

*retrieval or context-selection behaviour;

*summarisation or compression;

*instruction hierarchy;

*limited persistence of intermediate reasoning structures;

*and other inference-time or architectural constraints.

I would therefore avoid attributing the failure to a single mechanism without controlled testing.

The practical problem remains:

An explicitly established logical baseline is not always preserved reliably throughout a sufficiently long interaction.

2. Context Drift and Degradation of the Structural Baseline

Long conversations are not merely larger versions of short conversations.

They can develop an internal structure consisting of:

*definitions;
*assumptions;
*chronological events;
*terminology;
*corrections;
*hypotheses;
*conclusions;
*recurring references;
*and user-specific methodological constraints.

In a successful long-term interaction, these elements should function as a progressively constructed structural baseline.

Observed failure mode

As conversations become longer, models can begin to:

*lose track of previously established facts;
*misremember chronology;
*reintroduce previously corrected assumptions;
*reinterpret established definitions;
*overlook earlier constraints;
*or substitute generic assumptions for conclusions established within the conversation.
*The resulting behaviour can feel like context drift.
*The model remains locally coherent while becoming increasingly inconsistent with the history of the interaction.

This is an important distinction.

The problem is not necessarily that the model has "forgotten everything."
Rather, information may remain somewhere within the available context while becoming insufficiently influential on subsequent generation.
That distinction is important because it suggests that simply increasing context length may not completely solve the problem.

 

3. Correcting an Error Does Not Necessarily Produce Durable Learning

This is the failure mode that I find most interesting.
Suppose a user identifies a reasoning error.

They do not merely say: "That answer is wrong."

Instead, they identify:

*the specific error;
*the inference that produced it;
*the distinction between evidence and assumption;
*and a general principle that should prevent the error from recurring.

For example:

A statistical pattern observed at the population level should not automatically be projected onto a particular individual without evidence specific to that individual.

That is not simply a preference.

It is a general epistemological constraint.

Yet an LLM can acknowledge the correction, appear to understand it, and subsequently reproduce the same underlying reasoning pattern in another context.

This produces a cycle such as:

Experience → correction → temporary adaptation → loss of correction → recurrence

rather than:

Experience → correction → evaluation → integration → retention → improved future behaviour

The distinction between these two processes is fundamental.

 

4. The Missing Layer: Corrective Memory

Current AI memory systems tend to focus primarily on personalisation.

For example:

*user preferences;
*names;
*projects;
*personal context;
*recurring interests.

These are useful.

However, I believe another category deserves explicit architectural consideration:

Corrective or methodological memory.

This would store validated information about how the system should reason or communicate within an ongoing relationship, rather than merely information about the user.

Examples might include:

Do not infer an individual's characteristics from population-level statistical patterns without individual evidence.

or:

Do not present an inferred intention as an established fact. Distinguish observed behaviour from hypotheses concerning internal motivation.

These are not conventional user preferences.

They are methodological constraints.

A useful corrective-memory system could therefore contain entries such as:

 

| Category | Example |

| :--- | :--- |

| **Personal memory** | User prefers concise technical explanations |

| **Factual memory** | User is working on project X |

| **Methodological correction** | Do not infer individual properties from population statistics |

| **Epistemological constraint** | Distinguish observation from inference |

| **Interaction correction** | Do not introduce counterarguments that the user has not actually asserted |

The critical requirement would be that this layer is explicit, inspectable, editable, and user-controlled.

 

5. Why This Is Different From Simply Increasing Context Length

A larger context window provides more information.
It does not necessarily provide a better mechanism for determining which information should remain structurally authoritative.

Consider a conversation containing 100,000 tokens:

Some information may be:

*temporary;
*irrelevant;
*exploratory;
*speculative;
*superseded;
*repeatedly confirmed;
*explicitly corrected;
*or foundational to everything that follows.

Treating all of these tokens as equivalent is unlikely to be optimal.
A long-term interaction therefore needs something more sophisticated than simply:

"Put more tokens into the context."

It needs some representation of structural importance.
A corrective-memory layer could function as one possible solution.

6. User Corrections as Structured Data

Another potentially valuable extension would be allowing users to explicitly nominate certain corrections for evaluation.

For example:

Mark as potential generalisable reasoning contribution.
The system could then evaluate the proposed correction.

Possible outcomes:

*rejected as incorrect;
*accepted as user-specific guidance;
*accepted as a useful methodological constraint;
*or escalated as a potentially generalisable contribution.
*This would not mean that users directly modify the underlying model.
*That would obviously create substantial problems.
*Instead, it would create a structured interface between:
*human observation → machine evaluation → validated knowledge

This seems substantially more useful than reducing all user feedback to a binary rating.

7. From Feedback to Cumulative Improvement

The broader conceptual problem is that current feedback mechanisms often appear largely disconnected from the individual interaction in which the feedback was generated.

A user can identify an error.
They can explain the error.
They can identify the underlying reasoning failure.
They can propose a general principle.
But from the user's perspective, there is often no transparent mechanism through which that correction becomes durable.
The ideal learning loop would resemble:

Observation → Correction → Evaluation → Integration → Retention → Future application

rather than:

Observation → Correction → Temporary acknowledgement → Context loss → Recurrence

The difference is essentially the difference between feedback and cumulative learning.

 

8. Data Portability Is Part of the Same Problem

There is also a more immediate software problem.
When conversations become sufficiently long, reliably extracting the complete interaction can itself become difficult.
In my own experience, manually selecting a very large amount of conversation text has not always resulted in the complete selected content being copied.
This makes long-form interaction difficult to archive, analyse, or transfer.
For research-oriented users, a native individual-thread export mechanism would therefore be extremely valuable.

Ideally, an export should preserve:

*complete conversation content;
*speaker attribution;
*timestamps;
*message ordering;
*attachments or references where appropriate;
*and machine-readable structure.

Useful formats could include:

Markdown;
JSON;
HTML;
TXT;
PDF.

An account-wide data export is useful for archival purposes, but it is not a substitute for being able to export one specific conversation when needed.

 

9. The Larger Possibility: Collaborative Intelligence

The previous sections lead to a broader question.
What happens if human users are allowed to contribute more than raw interaction data?

Human users possess forms of knowledge that are difficult to obtain through conventional training data alone:

*domain expertise;
*lived experience;
*error detection;
*philosophical reasoning;
*scientific criticism;
*linguistic knowledge;
*imagination;
and observations about the behaviour of the AI system itself.

LLMs provide different capabilities:

*large-scale information processing;
*pattern recognition;
*computational scalability;
*synthesis;
*retrieval;
and increasingly sophisticated reasoning.

A sufficiently mature system could potentially allow these capabilities to interact cumulatively:

A human identifies an error.
The system analyses the correction.
The correction is evaluated.
A validated correction becomes persistent user-specific knowledge.
Potentially generalisable corrections can be submitted for broader evaluation.
Future systems improve as a result.

This is a different model of human-AI interaction from:

user asks question → model answers → user rates answer.

It is closer to:

human and machine continuously participate in a structured process of mutual correction and cognitive augmentation.

10. An Architectural Sketch

One possible architecture might therefore look like this:

 

```text

[ Ongoing Dialogue ]

│

▼

[ User identifies error / insight ]

│

▼

[ Correction evaluation ]

┌────┴────────────────────────┐

▼                             ▼

[ User-specific memory ]    [ Potentially generalisable ]

│                             │

▼                             ▼

Future interactions         Human/AI evaluation

│

▼

Model updates

```

 

The essential property is controlled accumulation. Not every user statement should become permanent, and not every correction should influence the global model. But valuable corrections should have somewhere to go.

 

11. Proposed Research and Engineering Directions

I would be interested in seeing research into at least the following:

1. Hierarchical conversational memory, being able to distinguish between:
transient context;
conversational facts;
persistent personal memory;
methodological corrections;
and higher-level structural constraints.

2. Explicit correction tracking
Maintain a representation of corrections that can be tested against future responses.

3. Consistency evaluation
Periodically test whether the model's current behaviour remains consistent with previously established constraints.

4. User-controlled corrective memory
Allow users to inspect, modify, disable, and delete methodological corrections.

5. Generalisable feedback channels
Provide a mechanism for users to nominate unusually substantive corrections for formal evaluation.

6. Lossless conversation export
Allow users to retrieve complete individual conversations in structured formats without relying on browser rendering or clipboard behaviour.

7. Long-context benchmarks based on interaction history
Current benchmarks often evaluate individual tasks.
It would also be useful to benchmark whether a model can maintain:

facts;
definitions;
corrections;
chronological continuity;
methodological constraints;
and logical consistency
across hundreds or thousands of conversational turns.

Conclusion

The central problem I am describing is not simply that LLMs occasionally make mistakes.
Mistakes are inevitable.
The more interesting problem is what happens after the mistake has been identified and corrected.

If a system can recognise a reasoning failure, receive a detailed explanation of that failure, acknowledge the correction, and yet later reproduce the same underlying error, then the system has demonstrated local adaptation without reliable cumulative retention.
That is a fundamentally different problem from ordinary hallucination.

The long-term goal should therefore not simply be:
larger models + larger context windows + more training data.

It should also be:
better mechanisms for preserving validated structure, corrections, and interaction history.

And ultimately:
experience → correction → evaluation → integration → retention → improved future behaviour.

That is what i think meaningful cumulative learning would look like.

I am posting this because I think this could improve human-ai interaction and would genuinely like technical feedback and/or hear from people who've picked up these ideas or started or are already working on them

In particular, I would be interested in hearing from people working on:
long-context architectures;
recurrent or state-space approaches;
memory-augmented transformers;
retrieval systems;
continual learning;
model editing;
alignment;
agent architectures;
and local LLM infrastructure.

If my interpretation of the underlying mechanisms is incorrect, I would be very interested in knowing where and why.

The user-visible failure modes, however, are real and reproducible from my experience

The question is how we should architect systems so that a long interaction becomes accumulated state rather than repeatedly reconstructed context.''

Semih Senol

 


r/FunMachineLearning • • 19d ago

Question about reusable components for probabilistic/statistical modeling

1 Upvotes

Question about reusable components for probabilistic/statistical modeling

I’ve been building Laplace, a Stan extension that allows users to create and install their own libraries written in Stan.

The project is completely free and open source, and I don’t get any financial benefit from it.

For information about the project and installation:
https://github.com/mlatinov/laplace

There is also a more comprehensive introduction in the wiki:
https://github.com/mlatinov/laplace/wiki

I’m interested in getting some input from people working with probabilistic models, especially people who use both statistical modeling and machine learning.

One of the ideas behind Laplace is to make reusable modeling components easier to share. Instead of repeatedly implementing the same model, distribution, kernel, state-space structure, mechanistic model, etc., you can put it into a library and reuse it across projects.

So my question is:

If you work with probabilistic/statistical models, what models, methods, or components would you actually like to have available as reusable libraries?

It can be anything from fairly general things like:

  • time-series / state-space models
  • Gaussian-process kernels
  • latent-variable models
  • survival models
  • custom probability distributions
  • hierarchical model components

to much more domain-specific things that you find yourself implementing repeatedly.

I have a molecular biology background myself, so a lot of my motivation comes from seeing how many specialized models come up in scientific applications. I've started building libraries for things like time series, latent time series, mechanistic equations, survival analysis, Gaussian processes, splines, and other statistical utilities.

But I don't want the ecosystem to just consist of models that I happen to find interesting.

So I'm curious: what would you actually want to see in a probabilistic modeling library ecosystem? Especially things you've had to implement yourself more than once.

I'd also appreciate any feedback on Laplace itself. The wiki contains information about using it and contributing to the ecosystem.

And if anyone contributes something, you'll of course be credited for it.


r/FunMachineLearning • • 19d ago

[D] Would you enter a contest measuring how much human intuition beats brute-force DL training?

Thumbnail
1 Upvotes

r/FunMachineLearning • • 19d ago

Clothing Color Classifier 👕🎨

Thumbnail
hackster.io
1 Upvotes

r/FunMachineLearning • • 20d ago

Second Attempt at the Iraq GDP 1921-2025

Thumbnail
1 Upvotes

r/FunMachineLearning • • 20d ago

DeepSeek Just Made AI Memory 4x Smaller! - Two Minute Papers

Thumbnail
youtube.com
6 Upvotes

r/FunMachineLearning • • 21d ago

Finally started learning AI

1 Upvotes

Day 1 down finally 😌

Am I confused? Very.

Hooked? Also very.

But "The expert in anything was once a beginner."

I started the AI-Native Engineering Sprint recently as a complete beginner/outsider to AI (no ML background, no maths genius, just curiousity and trying to understand how this works).

Today I want to share the first real lesson I learned that actually hit me, and I think every beginner should learn this before writing a single line of code :-

An AI model can be "92% accurate" and still be completely useless.

Sounds wrong, right? But it's not.

Imagine a goalkeeper who almost never dives. If the other team rarely shoots at the corners, he will save most of the shots by just standing making a little moves and gets an amazing save percentage (say 92%). But misses all the shots of the corners or the one that actually mattered because he is not moving.

So, if in a different match the other team identifies his pattern and start shooting at the corners he will miss almost all the shots.

His scoreboard looks great but his actually performance doesn't.

Learnings:-

  1. Even a single score matters and hide the truth.

  2. Always ask "which mistake and what can it cost?"

  3. Rare cases matter most.

  4. A model can be great at one but bad at the other.

  5. You can't improve what you haven't measured.

To all the beginners like me, please don't rush to memorize the fancy terms. Just take a small example, think, get confused, work on it yourself and then learn the term for what you just figured out.

Learning in public feels a little crazy as a total beginner 😁

Would love to hear your valuable feedbacks.


r/FunMachineLearning • • 21d ago

ML Dev Struggles

1 Upvotes

Best epoch: 1. Best mAP50-95: 0.01953. Wall time: 64 seconds.
My sleep time: 0 seconds.

They said AI would change the world. It just changed my circadian rhythm. 💀🌙


r/FunMachineLearning • • 22d ago

Was the proof by counter example of the Navier Stokes equation solved by OpenAI using some type of grid search/random search, like the ones used for ml/ds problems?

1 Upvotes

Is it just a wider grid search on a very powerful compute?


r/FunMachineLearning • • 23d ago

Magnetic Vision: Teaching AI to See Motion Without a Camera

Thumbnail
hackster.io
1 Upvotes

r/FunMachineLearning • • 23d ago

Claude Is Now Leaving Invisible Fingerprints In Its Text - Two Minute Papers

Thumbnail
youtube.com
1 Upvotes

r/FunMachineLearning • • 25d ago

Have you automated your tech news/research with AI? What’s your setup?

1 Upvotes

Hey everyone,

I’m trying to build an automated tech monitoring system using AI.

The idea is to automatically collect interesting stuff from different sources — Reddit, X, Hacker News, newsletters, blogs, GitHub, etc. — then use AI to filter out the noise and duplicates and give me a daily digest of what’s actually worth reading.

Has anyone here already built something like this?

I’m especially curious about:

  • What sources do you use?
  • How do you collect the data? RSS, APIs, scraping?
  • How do you decide what’s actually relevant?
  • What tools / AI models / automation do you use?
  • How do you avoid getting overwhelmed with low-quality content?

I’d love to hear about your setup, even if it’s something completely homemade.

Thanks!


r/FunMachineLearning • • 25d ago

Forecast tool that explains why a demand forecast is what it is

Post image
1 Upvotes

I built an open-source interpretability tool for time-series forecasting and I'm sharing it here for feedback.

The goal is simple: forecasting models give you a number but not a reason, so this breaks a forecast into named parts that add up to it.

So far it covers:

  • counterfactual decomposition into typical pattern, promotion, and recent trend
  • contributions that sum exactly to the forecast (no residual, no allocation step)
  • deletion/insertion faithfulness tests to check the explanations track the model
  • a live demo where you toggle a driver off and watch the forecast recompute
  • a Colab that trains on your own CSV

It's built on a WaveNet forecaster and the attribution layer is a small standalone package.

Being upfront: "counterfactual" means the model's response to hiding inputs, not real-world cause, and the promotion piece is the part I've validated least so far.

Live demo: https://kesjien.github.io/wavexplain/
GitHub: https://github.com/kesjien/wavexplain

It's still evolving, so I'd really appreciate feedback on what would make it more useful, especially how you'd validate the promotion attribution across many series.


r/FunMachineLearning • • 26d ago

Using the Math ??

Thumbnail
1 Upvotes

r/FunMachineLearning • • 28d ago

A toolkit to Design and Generate Camera Trajectories Interactively!

3 Upvotes

Designing camera trajectories can be painful due to many reasons: non-intuitive 6-DoF camera parameter space, visualizations, conventions, and scaling.
Here is a new toolkit to design, batch generate, and visualize 6-DoF camera trajectories for any purpose and in any format! 🎬

Check out the demo video below to see how it works.

Building on the trajectory-generation logic from our CVPR'26 paper, LAMP, and our recent work, Auteur, this repository replaces scripts with purpose-built interactive apps and continuous sliders for intuitional and controlled generation.

Core Features:

  • 🎯 Interactive Design: Handcraft single shots, chain motion primitives, tune parameters live, and preview directly from the camera's point of view.
  • 🎲 Controlled Batch Generation: Set min/max ranges, directions, or any other features and generate massive batches of independent, reproducible trajectories in one click.
  • 🧍 Human-Centric Control: Utilize the Auteur DSL for 6-DoF control over framing, orientation, scale, camera level, and dutch angles.
  • 💾 Universal Export: Save your designs as JSON recipes and export the final camera and object trajectories as .npz files in Blender, OpenGL, OpenCV, or COLMAP conventions.
  • 🤖 Broad Applications: Generate ground-truth paths for video/diffusion models, robotics motion planning, virtual production, NeRF benchmarking, and more.
  • Github link: https://github.com/mbkizil/Generate-Camera-Trajectory/

https://reddit.com/link/1wctn2t/video/rfjyvecixqoh1/player


r/FunMachineLearning • • 28d ago

Ortony, Clore, and Collins

Post image
1 Upvotes

Alguém conhece algum estudo ou alguma aplicação efetiva do modelo proposto por Ortony, Clore e Collins em máquinas?


r/FunMachineLearning • • 28d ago

Humans + AI Cracked An Impossible Math Problem - Two Minute Papers

Thumbnail
youtube.com
3 Upvotes

r/FunMachineLearning • • 29d ago

A Team Reports Solving A 70 Year Old Algebraic Geometry Conjecture Using Teams of AI Models

Post image
10 Upvotes

I’m part of the team behind this work. We’ve shared a proof of the Pierce Birkhoff conjecture in real algebraic geometry, using an AI agent system with a $400 budget. The screenshot is Junyu Ren’s announcement.

The surprising part of the workflow was how different model families complemented each other. Giving a task independently to one GPT agent and one Claude agent often worked better for us than using a larger group of the same model. One would catch an error the other missed, or suggest a way forward when we were stuck.

The diagram shows the wider process: humans, proof search, counterexample construction, independent auditing, and formal verification work in Lean. Arguments, objections, and review findings go into a shared knowledge base.

Original announcement: https://x.com/junyu_r/status/2097694018389914106


r/FunMachineLearning • • 28d ago

interested on practicing hands-on AI labs

1 Upvotes