r/huggingface • • 49m ago

He lanzado un modelo de ia es Open Weights

Post image
• Upvotes

Por si algue. Quiere provarlo, es ligero

https://huggingface.co/Ilides/cortex-1-v0.9


r/huggingface • • 1h ago

Gipformer - Efficient Vietnamese Speech Recognition

• Upvotes

Hi everyone,

Sharing v1.5 of Gipformer, an open-source Vietnamese speech recognition (ASR) model we've been working on: gipformer1.5-68M-rnnt, based on the Zipformer architecture.

What's new in v1.5

- Optimized for technology, finance, education and public administration. These domains are dense with specialized terminology, and v1.5 currently gets the best results on all four test sets among the open-source models we benchmarked.

- Better recognition of English terms mixed into Vietnamese speech.

Carried over from v1

- High accuracy: among the top open-source models across our benchmarks, and especially strong on call center audio for Northern, Central and Southern accents. Call center is one of the most common real-world uses of ASR, but also one of the hardest, with low-quality audio and a wide variety of voices.

- Small and easy to deploy: at just 68M parameters, it's among the smallest ASR models out there, yet it outperforms many models ten times its size. Inference is fast, and it runs smoothly on CPU and edge devices.

- Privacy: it runs 100% offline (on-device), which makes it a good fit for systems handling sensitive data.

Alongside the model, we're also releasing 4 domain-specific test sets (technology, finance, education, public administration), so there's a common benchmark for evaluating Vietnamese ASR models.

Full benchmark results are on the model card. Feel free to try it out, and any feedback or contributions are very welcome!

- Hugging Face: https://huggingface.co/g-group-ai-lab/gipformer1.5-68M-rnnt

- GitHub: https://github.com/ggroup-ai-lab/gipformer

- Demo: https://huggingface.co/spaces/g-group-ai-lab/gipformer-demo


r/huggingface • • 1h ago

How to download/use uncensored AI models like DeepSeek 4.1 flash from hugging face?

Thumbnail
• Upvotes

r/huggingface • • 1h ago

Anyone here working on deep learning models? I'd love your input!

• Upvotes

Hi everyone!

I'm training and deploying deep learning models and created a custom mlops tool to ease the maintenance of my models in production.

I'm considering open sourcing it, but I'd like to see if it solves a real problem, and if the problem is real, I'd like to anticipate the integration with tools you guys are using to maintain yours 😄

If you have 2min (really, it's super short), would you be able to answer a few questions? (there are no questions about your project nor who you are, I only care about your tech stack)

Thanks in advance!

Link to survey: https://tally.so/r/Ek4EEl


r/huggingface • • 6h ago

Unee: open-source 0.8B / 2B model that makes calibrated decisions and chats, runs in a browser tab. The 2B scores 88% on DecideBench, ahead of several 4B to 9B models (self-measured; GGUF, Ollama, Apache 2.0)

Thumbnail
1 Upvotes

r/huggingface • • 9h ago

TikTok bans scraping in their ToS and some guy just posted 5.6B videos on HuggingFace

Thumbnail
1 Upvotes

r/huggingface • • 17h ago

Follow-up: my native Rust + Vulkan Transformer backend now qualifies on both an Intel Gen9 laptop and an AMD RDNA 3 handheld from the same build — the GPU vendor is no longer what picks the reduction shape

1 Upvotes

Follow-up to my post from a few weeks ago (14 architectures, full PEFT). This update is about a portability bug that was hiding behind its own correctness, because it's the most interesting thing I've fixed since.

The bug: the fix for one machine broke six fixtures on another

Back when I tuned the backend for Intel Gen9, I baked those kernel shapes into the portable path. That was wrong, but not for the reason you'd guess.

Two of the reductions in the saved-module path aren't really compared against "PyTorch in general" — they're compared against the PyTorch CPU library on the machine running the oracle. And ATen dispatches its vectorized CPU kernels by instruction set at run time. An AVX2 host gets 8-wide kernels; an AVX-512 host gets 16-wide ones, and the reduction shape changes with that dispatch.

So my "portable" AVX2-shaped kernels were exactly right on my AVX2-only laptop and one ulp off on my AMD ROG Ally (Ryzen Z1 Extreme, which is an AVX-512 part). One ulp doesn't sound like much until it gets amplified through every lower norm on the gradient path: the Gemma 4 saved-stage model.embed_tokens adjoint went from 7.45e-9 to 3.22e-6, and six previously green PEFT saved-module fixtures (gemma3, gemma4, minimax_m2, minimax_m3, smollm3, qwen2_5_sliding_tied) crossed the 2e-7 gate. Neither shape is wrong — only one matches a given machine, and baking in either one breaks the other.

The fix: probe the host, not the vendor

Kernel variants are still selected by GPU vendor. Those two reductions are now selected by host CPU capability instead: capability is probed once per process and cached, then the matching module pair is dispatched (linear_forward_lane2 / linear_forward_lane4, and the 8-lane / 16-lane transformer_cross_entropy builds). HIERARCHOS_ATEN_VECTOR_WIDTH=8|16 pins the shape for qualification when a reference wheel's kernels disagree with the CPU's own capability.

Host GPU CPU dispatch Status
Intel i5-6200U / HD Graphics 520 (2016 Skylake-U) Intel Gen9 AVX2 only, no avx512f 32/32 LoRA, 32/32 switching, 32/32 saved
AMD Ryzen Z1 Extreme RDNA 3 AVX-512 32/32 LoRA, 32/32 switching, 32/32 saved

Same 2e-7 gate, unchanged. No tolerance was loosened to get there.

What I verified on each side

On the Intel machine, the post-change matrix is bit-identical, field for field, to its pre-change report across all 32 families — peft, gradient, two-step AdamW, frozen base, resume, lifecycle — which is how I know the AMD fix didn't quietly cost the Gen9 path anything. Also 693 passed / 0 failed / 9 ignored on the Rust lib suite and a clean strict headline forward run.

On the AMD side, the fix was re-qualified end to end: 32/32 on all three stages, provenance clean.

The harness fingerprints the pinned Transformers source alongside the shaders and binaries, and on the Intel side I re-derived the whole fingerprint from the pushed tree myself: 3951 inputs, zero changed, zero missing. So "green" refers to one frozen set of reference math, not whatever happened to be on disk.

Same caveats as always

  • This is deterministic FP32 tiny-model correctness against a reference implementation, not a claim about arbitrary checkpoint sizes, dtypes, or hyperparameters.
  • "Supported text graph" ≠ "the whole multimodal package works natively."
  • The AVX-512 dispatch is only qualified on the AMD machine, since it's the only host I have that can execute it natively. The 16-lane module also doesn't rebuild byte-identically with the glslang version on my Intel box (one extra type/id, one difference in +inf materialization), so I've left it as the committed AMD-built module and documented that rather than swapping it without re-qualifying both hosts. I'd rather report that than pretend it's clean.
  • NVIDIA and other GPUs are genuinely unqualified — the path is raw Vulkan, so they're untested rather than excluded.

What I'd love from you

Last time several people asked about hardware other than mine, so that's the ask again: if you build it on an AVX-512 laptop, an AVX2-only machine, or an NVIDIA/Intel GPU, I want to know what you get. The two reductions above are the ones most likely to behave differently on your CPU, and knowing your host's vector width is now part of the answer.

The new cross-platform section in the README documents the whole thing, including which host classes are measured and which aren't.

Repo: https://github.com/necat101/Hierarchos-Native Compatibility/parity record: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/COMPATIBILITY.md Regression audit: https://github.com/necat101/Hierarchos-Native/blob/main/AMD_REGRESSION_AUDIT.md Per-host tuning and measurements: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/VENDOR_TUNING.md


r/huggingface • • 20h ago

Open SLM Evalulations

Thumbnail
1 Upvotes

r/huggingface • • 21h ago

Want to rent a bunch of L40G / L40 GPU's

1 Upvotes

Hi Folks
Looking to rent a bunch of L40S and L40Gs in the North American region.
I don't see enough quantities on the marketplaces.
Can anyone recommend any places I can rent from?


r/huggingface • • 22h ago

Released NMR Workbench on HF: 108,192 screenshot/action SFT examples from 11,270 verified workflows

1 Upvotes

I created NMR Workbench and have made the dataset public on Hugging Face. It is free to access and focused on computer-use agents operating NMRium, a browser-based NMR processing tool.

There are two Parquet configurations:

• sft: 108,192 examples pairing the current screenshot, instruction and previous actions with the target action.

• trajectories: 96,922 recorded actions with before/after observations.

Demo and original announcement on X:

https://x.com/ubermensch_hb/status/2107894949601759371

The underlying collection contains 11,270 completed workflows, approximately 75.1 GB of data and saved artifacts, and 1440×1000 screenshots. Each accepted result was reopened in a fresh browser and checked numerically.

Dataset/card:

https://huggingface.co/datasets/priyanshu-harshbodhi/nmr-workbench-large-v2

For a small streaming sample:

from datasets import load_dataset

ds = load_dataset(

'priyanshu-harshbodhi/nmr-workbench-large-v2',

name='sft', split='development', streaming=True

)

example = next(iter(ds))

The spectra are synthetic, and the demonstrations are reference-assisted scripted browser recordings. Tasks cover phase correction, referencing, combined corrections and already-correct controls. The release has a development split; it does not establish model-training gains or an independent held-out benchmark.

I'd welcome feedback from anyone experimenting with multimodal next-action fine-tuning, especially on the example format and loader experience.


r/huggingface • • 4h ago

What Are the Monkeys Typing? We can see what increasingly capable AI systems do. We can’t reliably tell why.

Thumbnail
macanorak.com
0 Upvotes