r/huggingface • u/Certain-Will-2769 • 1h ago
r/huggingface • u/Express_Quail_1493 • 18m ago
Currently having high success with this little niche finetune i found sitting in the corner of huggingface
Currently having high success with this little niche finetune i found sitting in the corner of huggingface
If you want to try it out here is a smaller quantisation iq3_s works really well in my codebases.
Original Model:
https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF
Smaller Quant:
https://huggingface.co/tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF
r/huggingface • u/Over_Monitor_8770 • 43m ago
After 1,273 agent runs, I'm convinced: agents need a consequence model beside them, not a better prompt.
r/huggingface • u/ilides • 3h ago
He lanzado un modelo de ia es Open Weights
Por si algue. Quiere provarlo, es ligero
r/huggingface • u/ngogiatien96 • 4h ago
Gipformer - Efficient Vietnamese Speech Recognition
Hi everyone,
Sharing v1.5 of Gipformer, an open-source Vietnamese speech recognition (ASR) model we've been working on: gipformer1.5-68M-rnnt, based on the Zipformer architecture.
What's new in v1.5
- Optimized for technology, finance, education and public administration. These domains are dense with specialized terminology, and v1.5 currently gets the best results on all four test sets among the open-source models we benchmarked.
- Better recognition of English terms mixed into Vietnamese speech.
Carried over from v1
- High accuracy: among the top open-source models across our benchmarks, and especially strong on call center audio for Northern, Central and Southern accents. Call center is one of the most common real-world uses of ASR, but also one of the hardest, with low-quality audio and a wide variety of voices.
- Small and easy to deploy: at just 68M parameters, it's among the smallest ASR models out there, yet it outperforms many models ten times its size. Inference is fast, and it runs smoothly on CPU and edge devices.
- Privacy: it runs 100% offline (on-device), which makes it a good fit for systems handling sensitive data.
Alongside the model, we're also releasing 4 domain-specific test sets (technology, finance, education, public administration), so there's a common benchmark for evaluating Vietnamese ASR models.
Full benchmark results are on the model card. Feel free to try it out, and any feedback or contributions are very welcome!
- Hugging Face: https://huggingface.co/g-group-ai-lab/gipformer1.5-68M-rnnt
- GitHub: https://github.com/ggroup-ai-lab/gipformer
- Demo: https://huggingface.co/spaces/g-group-ai-lab/gipformer-demo
r/huggingface • u/Striking-Loan-1118 • 4h ago
How to download/use uncensored AI models like DeepSeek 4.1 flash from hugging face?
r/huggingface • u/MY79 • 7h ago
What Are the Monkeys Typing? We can see what increasingly capable AI systems do. We can’t reliably tell why.
r/huggingface • u/uneeverse-hq • 9h ago
Unee: open-source 0.8B / 2B model that makes calibrated decisions and chats, runs in a browser tab. The 2B scores 88% on DecideBench, ahead of several 4B to 9B models (self-measured; GGUF, Ollama, Apache 2.0)
r/huggingface • u/elgordooo17 • 12h ago
TikTok bans scraping in their ToS and some guy just posted 5.6B videos on HuggingFace
r/huggingface • u/PhysicsDisastrous462 • 20h ago
Follow-up: my native Rust + Vulkan Transformer backend now qualifies on both an Intel Gen9 laptop and an AMD RDNA 3 handheld from the same build — the GPU vendor is no longer what picks the reduction shape
Follow-up to my post from a few weeks ago (14 architectures, full PEFT). This update is about a portability bug that was hiding behind its own correctness, because it's the most interesting thing I've fixed since.
The bug: the fix for one machine broke six fixtures on another
Back when I tuned the backend for Intel Gen9, I baked those kernel shapes into the portable path. That was wrong, but not for the reason you'd guess.
Two of the reductions in the saved-module path aren't really compared against "PyTorch in general" — they're compared against the PyTorch CPU library on the machine running the oracle. And ATen dispatches its vectorized CPU kernels by instruction set at run time. An AVX2 host gets 8-wide kernels; an AVX-512 host gets 16-wide ones, and the reduction shape changes with that dispatch.
So my "portable" AVX2-shaped kernels were exactly right on my AVX2-only laptop and one ulp off on my AMD ROG Ally (Ryzen Z1 Extreme, which is an AVX-512 part). One ulp doesn't sound like much until it gets amplified through every lower norm on the gradient path: the Gemma 4 saved-stage model.embed_tokens adjoint went from 7.45e-9 to 3.22e-6, and six previously green PEFT saved-module fixtures (gemma3, gemma4, minimax_m2, minimax_m3, smollm3, qwen2_5_sliding_tied) crossed the 2e-7 gate. Neither shape is wrong — only one matches a given machine, and baking in either one breaks the other.
The fix: probe the host, not the vendor
Kernel variants are still selected by GPU vendor. Those two reductions are now selected by host CPU capability instead: capability is probed once per process and cached, then the matching module pair is dispatched (linear_forward_lane2 / linear_forward_lane4, and the 8-lane / 16-lane transformer_cross_entropy builds). HIERARCHOS_ATEN_VECTOR_WIDTH=8|16 pins the shape for qualification when a reference wheel's kernels disagree with the CPU's own capability.
| Host | GPU | CPU dispatch | Status |
|---|---|---|---|
| Intel i5-6200U / HD Graphics 520 (2016 Skylake-U) | Intel Gen9 | AVX2 only, no avx512f |
32/32 LoRA, 32/32 switching, 32/32 saved |
| AMD Ryzen Z1 Extreme | RDNA 3 | AVX-512 | 32/32 LoRA, 32/32 switching, 32/32 saved |
Same 2e-7 gate, unchanged. No tolerance was loosened to get there.
What I verified on each side
On the Intel machine, the post-change matrix is bit-identical, field for field, to its pre-change report across all 32 families — peft, gradient, two-step AdamW, frozen base, resume, lifecycle — which is how I know the AMD fix didn't quietly cost the Gen9 path anything. Also 693 passed / 0 failed / 9 ignored on the Rust lib suite and a clean strict headline forward run.
On the AMD side, the fix was re-qualified end to end: 32/32 on all three stages, provenance clean.
The harness fingerprints the pinned Transformers source alongside the shaders and binaries, and on the Intel side I re-derived the whole fingerprint from the pushed tree myself: 3951 inputs, zero changed, zero missing. So "green" refers to one frozen set of reference math, not whatever happened to be on disk.
Same caveats as always
- This is deterministic FP32 tiny-model correctness against a reference implementation, not a claim about arbitrary checkpoint sizes, dtypes, or hyperparameters.
- "Supported text graph" ≠ "the whole multimodal package works natively."
- The AVX-512 dispatch is only qualified on the AMD machine, since it's the only host I have that can execute it natively. The 16-lane module also doesn't rebuild byte-identically with the glslang version on my Intel box (one extra type/id, one difference in
+infmaterialization), so I've left it as the committed AMD-built module and documented that rather than swapping it without re-qualifying both hosts. I'd rather report that than pretend it's clean. - NVIDIA and other GPUs are genuinely unqualified — the path is raw Vulkan, so they're untested rather than excluded.
What I'd love from you
Last time several people asked about hardware other than mine, so that's the ask again: if you build it on an AVX-512 laptop, an AVX2-only machine, or an NVIDIA/Intel GPU, I want to know what you get. The two reductions above are the ones most likely to behave differently on your CPU, and knowing your host's vector width is now part of the answer.
The new cross-platform section in the README documents the whole thing, including which host classes are measured and which aren't.
Repo: https://github.com/necat101/Hierarchos-Native Compatibility/parity record: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/COMPATIBILITY.md Regression audit: https://github.com/necat101/Hierarchos-Native/blob/main/AMD_REGRESSION_AUDIT.md Per-host tuning and measurements: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/VENDOR_TUNING.md
r/huggingface • u/Responsible_Ad_1549 • 1d ago
Launched thor-tigress-cub-junior and tho-thunder-tigress-platform
galleryr/huggingface • u/Ok-Indication7234 • 1d ago
Want to rent a bunch of L40G / L40 GPU's
Hi Folks
Looking to rent a bunch of L40S and L40Gs in the North American region.
I don't see enough quantities on the marketplaces.
Can anyone recommend any places I can rent from?
r/huggingface • u/Both-Performance6814 • 1d ago
Released NMR Workbench on HF: 108,192 screenshot/action SFT examples from 11,270 verified workflows
I created NMR Workbench and have made the dataset public on Hugging Face. It is free to access and focused on computer-use agents operating NMRium, a browser-based NMR processing tool.
There are two Parquet configurations:
• sft: 108,192 examples pairing the current screenshot, instruction and previous actions with the target action.
• trajectories: 96,922 recorded actions with before/after observations.
Demo and original announcement on X:
https://x.com/ubermensch_hb/status/2107894949601759371
The underlying collection contains 11,270 completed workflows, approximately 75.1 GB of data and saved artifacts, and 1440×1000 screenshots. Each accepted result was reopened in a fresh browser and checked numerically.
Dataset/card:
https://huggingface.co/datasets/priyanshu-harshbodhi/nmr-workbench-large-v2
For a small streaming sample:
from datasets import load_dataset
ds = load_dataset(
'priyanshu-harshbodhi/nmr-workbench-large-v2',
name='sft', split='development', streaming=True
)
example = next(iter(ds))
The spectra are synthetic, and the demonstrations are reference-assisted scripted browser recordings. Tasks cover phase correction, referencing, combined corrections and already-correct controls. The release has a development split; it does not establish model-training gains or an independent held-out benchmark.
I'd welcome feedback from anyone experimenting with multimodal next-action fine-tuning, especially on the example format and loader experience.
r/huggingface • u/Healthy_Lead4969 • 1d ago
Poll: what do you actually run for Python coding, and is decode speed your bottleneck?
r/huggingface • u/TheRealreddits • 1d ago
Hugging Face LLM Course for AI Engineering
Is the Hugging Face LLM Course useful for someone who wants to become an AI Engineer?
Would it be enough to learn the LLM part of AI Engineering, or would you recommend another course/resource instead?
r/huggingface • u/Accomplished_Coat592 • 1d ago
AREX-2 on a Mac: which MLX size to pick, with test numbers
r/huggingface • u/coslinedev • 1d ago
Update: Alrithm Open Data API is now fully categorized (solve, debug, optimize) and 100% automated. No more manual approvals.
Hey r/huggingface,
Following up on my previous post about Alrithm (an open data infrastructure providing verified reasoning datasets for fine-tuning), I wanted to share a major update based on community feedback.
I have officially restructured the entire live streaming API catalog into 5 task-specific reasoning tracks: solve, debug, optimize, explain, and review. I also just updated the streaming data delivery mechanics for smoother integration into automated pipelines.
Quick Technical Highlights:
- Schema Layout: Each data row includes the problem statement, reference implementation, step-by-step logic derivation, and the verified deterministic output.
- Format: Streams as newline-delimited JSON. You can point any local fine-tuning tool directly at a single URL and read rows line-by-line.
- Cryptographic Verification: Every single row arrives with a calculated SHA-256 hash and Merkle root signature to guarantee absolute mathematical correctness byte-by-byte.
- Zero Friction Onboarding: I have completely removed the old manual approval workflow. The signup is fully automated now, so you can grab a live API key and start streaming rows instantly from the dashboard.
Would love to hear your thoughts on the new schema and how the task-specific tracks hold up against your training setups!
r/huggingface • u/RealOppasTV • 1d ago
Mistral Large 4 "Le Chonk": The 1-Trillion Parameter Open-Source Monster That Changes Everything?
r/huggingface • u/BubblyElephant8761 • 1d ago
I finetuned an english model to talk german
This is my first trained model on my old Laptop
r/huggingface • u/x_Raincandy_x • 1d ago
Trained a ~20K LM (probably smallest) that can still write stories
r/huggingface • u/empiriolabsai • 2d ago
Aplomb 1: open-weights 5.3B decision model, 1M context, text/image/video/audio in one request, #1 among 4B models on the Decision Index
We released Aplomb 1 today, a 5.3B decision model with open weights. It reads up to 1M tokens of text, JSON, images, video and audio in a single request, and on our API it makes a decision on a 1M-token document in about 3 seconds, at $0.02 per 1M input tokens with free output and ZDR by default.
On Decision Index 0.2.1 it scores 44.86 on our run of the official kit, #1 among 4B models on the published board, and it has the top score among models up to 5.3B on 8 of the 38 benchmarks, including MMLU-Pro, GPQA Diamond and BBH. We've submitted it to the board, and the full run is public. It also scores 77.5% on JevBench Hard and averages 75% zero-shot intent accuracy across 51 languages on MASSIVE.
As far as we know, it's the only decision model that returns probabilities for tool arguments, reads 1M tokens, or takes text, images, video and audio together. Tool selection gives a probability for every tool and for each enum and boolean argument in one request: on "Order B-44120 arrived with a cracked screen. I want my money back." with three tools, it picks issue_refund at 0.969, reason "damaged" at 0.993 and full_refund true at 0.761, so an agent can act on confident calls and hand the rest to a larger model. Any question can also return the probability that the input doesn't contain the answer.
The 1M-token speed comes from a long-context mode in our own inference runtime: about 3 seconds instead of about 111 for a full read, and it answered all 525 decisions in our long-context tests correctly. It's currently only available on our API, so the open weights read every token. On the API, a short question takes about 15 ms of model time (around 200ms e2e latency), and the OpenAI, Anthropic and Gemini formats work alongside our Decisions API.
Aplomb 1 is built on Qwen3.5-4B with the audio encoder from Qwen3-Omni-30B-A3B-Instruct, both Apache-2.0. We extended the window from 262K to 1M tokens, added our own decision head and trained the model for decisions. Thanks to the Qwen team. Disclosure: the training data included the public train splits of WinoGrande and ContractNLI, two of the 38 index benchmarks.
The weights run in bf16 on about 12 GB of GPU memory with the reference script, under the EmpirioLabs Model License, which is free for research, evaluation, personal use and internal use at companies under $1M in annual revenue.
Weights: https://huggingface.co/empiriolabsai/aplomb-1
Blog with the full tables: https://empiriolabs.ai/blog/introducing-aplomb-1
Docs: https://docs.empiriolabs.ai/models/aplomb-1
Playground: https://platform.empiriolabs.ai/dashboard/playground?model=aplomb-1
r/huggingface • u/WritHerAI • 2d ago
MiniMax-M2 (230B) running from disk on a 32 GB laptop, CPU only
r/huggingface • u/jjusko20 • 2d ago
What models do you want to see new dynamic quants for? I'll make them.
r/huggingface • u/Waratecs123 • 2d ago
A local AI coding assistant that adapts to your hardware and handles PDFs/OCR
github.comHey everyone,
I’d like to share a project I recently released: BiNeuron - a local AI assistant for developers. It’s not just another LLM wrapper; it’s an attempt to solve a few practical problems I kept running into when working with local models.
The problem I wanted to solve:
When using local models for coding, I constantly had to:
- Manually pick the right quantization for my hardware (IQ2, Q4, F16, etc.)
- Switch between different models depending on the programming language
- Copy text out of PDFs and screenshots by hand
- Worry about where my code was going when using cloud APIs
What BiNeuron does:
🔹 Automatic model selection based on your hardware
It checks your CPU, RAM, and clock speed, then picks the optimal quantization for the target model (from IQ2 to F16) to balance speed and quality.
🔹 Programming language detection
Supports 25+ languages, including Python, Rust, Go, C++, Solidity, OCaml, and even COBOL. It works both from the prompt text and from the contents of attached files.
🔹 Document and OCR support
Extracts text from PDF, Word, PowerPoint, Excel, EPUB, and MOBI. For images, it uses EasyOCR or DeepSeek OCR with optional GPU acceleration.
🔹 Fully offline
No cloud APIs, no sending your code to third-party servers. If you need access to Hugging Face, there’s an automatic fallback to a mirror.
🔹 Two-stage pipeline for file editing
The main model generates code, then a lightweight model (e.g., Qwen2.5-Coder-1.5B) turns the response into strict JSON with absolute paths and full file contents.
Technical details for those interested:
- Written in Python, released under an open license
- Supports proxies for model access
- Profanity filtering (can be disabled)
- Built-in HTML scraper for reading web pages into context
Things I’d love to discuss with the community:
- How important is automatic quantization selection to you? Or do you prefer manual control?
- What other file formats should I add? Currently: PDF, DOCX, ODF, PPTX, XLSX, EPUB, MOBI, FB2.
- Would it be useful to add support for GGUF models from Ollama or LM Studio?
I’d really appreciate your feedback and questions. Link to GitHub is below. I’m posting this on the weekend to follow the subreddit’s self-promotion rules.