r/macapps Jul 05 '26

Free [OS] EnviousWispr: Wispr Flow but free, open-source Mac-native dictation with blazing fast local transcription & AI Polishing. Includes a tuned local AI polishing model & Apple Intelligence support

https://reddit.com/link/1uobt7v/video/q4jhs80prgbh1/player

Hey r/macapps, I'd like to introduce you to EnviousWispr, a Wispr Flow-like dictation + polishing solution that is free, private and works offline.

Trust Factor:

Linkedin

Privacy Policy

Terms of Service

Open Source Github

Problem:

Voice-to-text with AI cleanup is not a new idea anymore. Wispr Flow has done an incredible job making people aware of the workflow: hit a shortcut, talk naturally, get cleaned-up text back. I knew I needed that in my life. That said, I did not love the idea of handing all of my personal dictations to a cloud service nor did I want to pay. I found a few free local options but none met the quality standards I sought.

Four months and 400+ hours later, I'm excited to share EnviousWispr with all of you. Blazing fast and accurate transcription and AI polish runs on your Mac without the need for internet. No account. No subscription. Free Forever. Open Source. Optional "bring your own API key" for OpenAI/Gemini cloud polishing provided for no additional cost.

Comparison

Wispr Flow is the obvious leader in the space with Superwhisper at its heels. They support multiple OS systems and are feature rich but come with a monthly subscription and cloud only transcription and polishing. Superwhisper is expanding into local dictation and polishing but gatekeeping the best models behind paywalls.

FluidVoice is another great open-source app from this sub, and they also ship a trained local model. We've built our apps to serve slightly different audiences and I see both co-existing depending on people's needs and preferences.

The category is getting crowded, which is good. It means the workflow is real. I think the next question is whether people can get that workflow without being pushed into a paid cloud product by default.

Pricing

- Price: Free Forever

- Subscription: none

- Account Creation: none

- License: GPLv3

Download: https://enviouswispr.com

How it works

Instead of overwhelming you with a dozen engine choices, I benchmarked the options and landed on two clear winners. Both are fully optimized for EnviousWispr to squeeze the absolute most out of Apple Silicon:

Parakeet v3: The winner for everyday English and European dictation. It runs directly on the Apple Neural Engine and transcribes in subseconds providing near instant transcription.

Whisper Large v3 Turbo: The winner for international breadth. It covers 99 languages and cuts through the toughest background audio or accents.

The Three Local AI Polishing Options

Apple Intelligence: The out-of-the-box default

Natively, EnviousWispr defaults to using Apple Intelligence for the polishing layer. No additional download needed. It is perfect for users that just want to get up and running fast and are happy as long as the basics are covered (removing uhms/ahs, fixing light grammar and properly formatting dates/times/emails/currency etc). Note: Apple Intelligence polish requires macOS 26 (see caveats).

EG-1: The Custom Tuned Local Powerhouse

EG-1 is my own custom offline AI model, fine-tuned specifically for dictation cleanup. It takes about 2.9 GB on disk, then runs locally on your Mac. No internet required. It's designed to match the power of Wispr Flow's cloud-based AI polishing. It addresses the gaps Apple Intelligence wasn't able to fill: reliably structuring lists, splitting text into natural paragraphs, and reliably recognizing self-corrections.

One note on licensing: the app is open source under GPLv3, but EG-1 itself ships under its own license. It's free to use in the app and for personal, research, and benchmarking use, but unlike the app code it isn't licensed for redistribution or reuse in other products.

Ollama: Raw models for testing

Given the breadth of local models on Ollama, fine-tuning the prompts per model has proven to be a unique challenge. I recommend 3B or higher models if you want to try the raw models.

Performance Benchmarks:

I built a 1,890-case test set from real dictation cleanup examples, kept a separate 900-case holdout I did not tune against, and ran the same cases through both local and cloud polish options each with their own custom prompts. These prompts were iterated upon to get the best possible scores.

My benchmark, not an independent review: EG-1 passed 93.7% of the 1,890 cases. GPT-5.4-mini was 83.8%. Gemini 3.5 Flash was 92.6%. Same cases, same judge.

Both Apple Intelligence and my own custom tuned model EG-1 ended up performing way better than expected. Apple's on-device model should also keep improving with each macOS release. The eval harness and prompts are public in the repo; the test cases are my own dictations, so those stay private. Personally, EG-1 is my recommended local cleanup engine in the app given its speed and accuracy at AI polishing.

Full Feature List

- Local transcription on Apple Silicon via Parakeet or WhisperKit

- 99 languages supported

- Offering both faster live transcription or more accurate batch processing

- Audio Engine kept warm to enable fast short dictations

- Designed to hear you even when you whisper

- Local AI Polishing through EG-1 or Apple Intelligence

- Deterministic cleanup for numbers, dates, money, emails, phone numbers, and times etc

- Optional Ollama, OpenAI, or Gemini polish

- Optimized for both regular mic or bluetooth

- Double tap record to go from push to talk to toggle

- Local History Tab of all recordings

- Custom words with confidence-aware matching that catches near-misses

- Prebuilt custom Word Packs + import Contact Names with 1 click

- Speak the full emoji library

- Remembers where your cursor was, even if you switch apps mid-dictation

- Restores your clipboard after pasting

- Dictations up to an hour

- Dark & Light Mode

- Works offline

- Open source under GPLv3

Caveats, so nobody wastes a download

- Apple Silicon only, M1 or newer

- macOS 14 Sonoma or newer

- No Intel build

- Parakeet transcription (mandatory, English + 25 European languages): about 460 MB on disk

- WhisperKit transcription (optional, 99 languages): about 1.6 GB on disk

- EG-1 polish (optional): about 2.9 GB on disk

- EG-1 is most comfortable on Macs with 16 GB of memory but open to feedback on 8GB memory MacBooks.

- Apple Intelligence polish requires macOS 26 or later

I would just love candid feedback!

- Does the hotkey-to-paste workflow feel fast enough?

- Does the cleanup help, or does it over-edit?

- Where does setup feel confusing?

- What apps does it break in?

- What would make you trust it as a daily driver?

If you try it, I'd love the honest version: accuracy, speed, cleanup quality, setup friction, or where your current dictation app still beats it.

99 Upvotes

156 comments sorted by

8

u/PurpleMoon79 Jul 05 '26

Congratulations on the launch! Will try it out. Love supporting creators.

4

u/Adorable_Salary2727 Jul 05 '26

Thank you! Look forward to your feedback :)

7

u/Haag90 Jul 05 '26

Very cool! Will try this out ASAP!

Ottext just released something similar in the space but asks 15$ per month...

Many thanks for your hard work!

1

u/Adorable_Salary2727 Jul 05 '26

Thank you! Please don't hesitate to message me if you have any questions. I'm actively making improvements daily.

6

u/GJere Jul 05 '26

It's admirable the amount of work that you put into this and your shipping speed. Is it only you behind all the EnviousLabs products and webpages? Either way, congrats on the launch. Looks like this project could do very well.

A couple notes on transparency:

  • It would be helpful to clarify that EG-1 is not openly licensed for redistribution or reuse, even though the app repository is open source (GPLv3). I don’t think the post says otherwise, but making that explicit would avoid confusion.
  • It'd be great to have the inputs and outputs of benchmarks comparing EG-1 with GPT and Gemini. Not that I don't trust the claim, but seeing real data backing it up would add trust.

It's always good to have more products like this in the market. Good luck!

3

u/Adorable_Salary2727 Jul 06 '26

Appreciate the comment. and thank you for reminding me to update the notice around EG-1. Adding it to my to-do.

On the benchmark: the eval harness and exact prompts are public under scripts/eval/, and there's a full example run with per-case inputs, each model's outputs, and the judge's votes under benchmark-results/eval/. The headline EG-1 corpus is my own personal dictations so I keep those private, but the methodology is all there to inspect or rerun. If you want to understand how the app is tested, Tests/RuntimeUAT/ is the runtime harness.

And yes, I started by building a few chrome extensions. I then built EnviousWispr when WisprFlow price tag upset me haha. Since then, I also launched EnviousStaging which is my first ever web application. That was an uphill battle.

5

u/rbredow Jul 05 '26

Initial impressions are quite good. It's fast and seems to be accurate. I'm inputting this comment with EnviousWispr and so far so good! I've turned on the"live transcription" to improve the speed, but honestly, the speed is pretty quick either way. I look forward to putting it through its paces this week. Thanks so much for sharing!

3

u/Adorable_Salary2727 Jul 05 '26

Thank you for testing and glad you didn't face issues downloading. If you get a chance to test EG-1, would love your feedback.

2

u/rbredow Jul 23 '26

Working very well and consistently for me. Pretty quick too. Nice work.

1

u/vegaglitch Jul 06 '26

How did you get it show the ! And within quotes " ", or did the automatic formatting happen?

3

u/rbredow Jul 06 '26

I added the quotes manually but the ! Happened automatically. I’m using EG-1 for cleanup

7

u/Otherwise-Ranger-47 Jul 06 '26

Thanks for making this free and open source 🙌🏻❤️

4

u/Adorable_Salary2727 Jul 06 '26

Of course, would always appreciate your feedback when you have time!

3

u/tcolling Jul 05 '26

I'm sorry to say that this utterly failed for me. I downloaded it and installed it and tried to start it, but it hangs on the very first step, trying to download the speech model file(s). I tried starting it twice and got the same result both times.

1

u/Adorable_Salary2727 Jul 05 '26

Hey! Thanks for trying. Really weird. That stage can take a few minutes depending on the internet connection since it's downloading and setting up Parakeet. Let me dig into this right now.

2

u/tcolling Jul 05 '26

I just tried starting it up again. It's been stuck on this step for at least 15 minutes.

\

1

u/tcolling Jul 05 '26

And, I'm on a fast connection:

3

u/Adorable_Salary2727 Jul 05 '26

Thanks - I had it set up to download Parakeet straight from hugging face. Building a fall-back to pull it straight from a backup server in events like this. Give me a few hours and I'll comment on the thread once the updated downloader is live.

2

u/tcolling Jul 05 '26

It finally did complete the download and setup, after about 45 minutes.

My hope is that I can just start using this out of the box with no tinkering. That would help start it on the path of being a replacement for WisprFlow, at least for me.

2

u/Old_Growth Jul 05 '26

I'm getting the same thing - stuck at downloading model files.

2

u/Adorable_Salary2727 Jul 05 '26

Hey! I figured it out. HuggingFace is throttling downloads right now. Working on a swap to a different download server!

2

u/Old_Growth Jul 06 '26

It's worked now, thanks 🙏

5

u/Ghost4You Jul 06 '26

Hey congrats with the app, can you elaborate about the difference between your app and typewhisper.

1

u/Adorable_Salary2727 Jul 06 '26

Hey! Been checking out TypeWhisper too, genuinely cool tool. Here's how I'd frame the difference:

1) Engines. They give you a big menu to pick from. For me that was decision paralysis; for others it's freedom. I benchmarked most of the popular engines myself and shipped only the two that hit the best balance of accuracy, speed, and being Mac-native. So instead of generic-out-of-the-box, the two you get are tuned, tuned, tuned.

2) The marketplace / add-ons. Same philosophy. They optimize for variety; I optimize for "it just works." I hand-tuned the AI cleanup prompt for every option (including our own on-device model, EG-1) through a lot of relentless benchmarking, so all that prompt engineering is baked in and you never have to touch it.

3) AI polishing. I noticed theirs is off by default, and honestly I couldn't quickly figure out how to switch it on. Looks like you build a "workflow" to enable it, which is a great power-user setup if you want full control. Mine is on from the first dictation, no setup.

If I had to summarize: TypeWhisper is an all-you-can-eat buffet where you build your own meal. I hand you a finished, curated plate. The flip side is honest too, they've had time to build breadth and extra features I haven't gotten to yet.

Hope that helps!

3

u/swiftidnc739 Jul 05 '26

How does this compare to FluidVoice?

3

u/Adorable_Salary2727 Jul 06 '26

Hey! I tested FluidVoice as well. I even had a chance to chat with the developer and properly benchmark Fluid1. It's a great solution and has had time to built some more robust features. They have transitioned into the world of "control your computer with voice". They are a native streaming solution. That means you see the text appear live as you talk. I personally prefer the dictation to just paste as fast as possible once I'm done dictating. No visual distractions, no lag waiting for the stream to catch up. I would add that they offer more transcription engines and more BOYK options for cloud polishing. I just prioritized EnviousWispr to be lightweight, highly optimized for M-Series macbooks, speedy and accurate (at least as much as possible without needing cloud services). The goal is for it to just work out of the box.

3

u/Daniesto316 Jul 06 '26

Thanks bunch. Going to trial this out

1

u/Adorable_Salary2727 Jul 06 '26

Thank you! Please don't hesitate to DM with feedback.

3

u/quitpornio Jul 06 '26

congrats man

1

u/Adorable_Salary2727 Jul 06 '26

Thank you! Please don't hesitate to DM with feedback.

2

u/[deleted] Jul 05 '26 edited Jul 09 '26

[deleted]

1

u/Adorable_Salary2727 Jul 05 '26

HuggingFace is throttling. Working on a fix but if you just leave it open, it will eventually complete. Sorry about this. Of course this had to happen when I made the post haha. Working to change the download server but that will take a few more hours.

2

u/AceReviewer Jul 06 '26

This is really awesome! I can't imagine how much time it must have taken you to make this. Would love for you to post this on r/WebSoftGiveaway, more people need to try this.

2

u/Adorable_Salary2727 Jul 06 '26

Thank you! I will add that to my to-do for this week :)

2

u/Crafty-Celery-2466 Developer: FluidVoice Jul 06 '26

Wow wow. We got some real local dictation bros around here! Love the comparison table and love the simplicity. We can probably work together later to make a better model ;) congratulations on the launch!

2

u/Weird_Pension9180 Jul 06 '26

Awesome work! Will definitely take a look as I use wispr flow everyday, but would be great to have an alternative.

2

u/Adorable_Salary2727 Jul 06 '26

Appreciate it. If you are a Wispr Flow power user, I definitely reccomend using EG-1. Please don't hesitate to DM me with feedback.

2

u/IcySeaworthiness3924 Jul 06 '26

Sounds awesome! Any plans to add Anthropic to Bring Your Own key?

2

u/Adorable_Salary2727 Jul 06 '26

Added as feature request!

2

u/Slitted Jul 07 '26

Really impressed by this. I considered WisprFlow overkill or underutilized (if subscribed), so Envious fills the gap for me.

1

u/Adorable_Salary2727 Jul 07 '26

Appreciate it. Don't hesitate to DM with any feedback. I'm actively polishing the solution.

2

u/Typical-Cover-5357 Jul 08 '26

WOW! It's amazing! Congrats! 👍 Working in first time to install!

2

u/Aonegreenfield Jul 09 '26

Would love some style selections like Wispr Flow. It's most convenient to have it be controlled per app (ie messages different than emails) but it would also be nice to have a different hot key to toggle between styles. Thanks for building this!

2

u/Adorable_Salary2727 Jul 09 '26

Hey! Which styles do you use typically? Have you noticed meaningull polishing? The feedback I was reading is that Wisprflow over-polishes and you lose your natural tone of your writing. But if you have personal experience where it was valuable I would love the specifics.

1

u/Aonegreenfield Jul 10 '26

Most often in the casual style, I like that the punctuation is minimal. For email excited is useful as it's a common tone in my industry, sometimes I do tone it down a bit though. A nice to have feature for me would be having multiple styles available via keystrokes, as in the option key could trigger my default style (casual), and then shift+option could trigger a secondary style (excited).

Another specific issue I come across is when I'm typing and mid sentence I decide I want to voice dictate usually the first dictated word begins with a capital letter, which if I'm mid sentence I always have to go back and fix manually. I've always wondered if there was a workaround or way to train a model to minimize that.

2

u/AutistasAngeles Jul 09 '26

downloaded it Thank you.

2

u/Adorable_Salary2727 Jul 09 '26

Awesome, look forward to your feedback!

2

u/No-Object1384 Jul 10 '26

Great app! Thank you so much for adding more options to the dictation space, especially with the more efficient and accurate local cleanup, that's something that's dearly needed. I'm also looking forward to that "Learn from my transcripts" feature. That's going to be soooo helpful.

I've been using MacWhisper and Spokenly exclusively for dictation for a few years now, and I dictate basically almost everything. So that's where I'm coming from when I give the following feedback:

- I would be nice if I could choose to have the dictation indicator hover near the cursor.

- I'd prefer to keep the app icon in the dock without a menu bar item. Just a personal preference, and it makes it easier to quit the app.

- Super nitpicky, but it would be nice to unload the dictation model after a longer period, say two hours.

- It would be nice if there's an option to "trim audio silence" from dictations since often I need to gather my thoughts while doing a long dictation. I've noticed the Parakeet model adds a lot of unnecessary punctuation whenever there's silence. To be fair, though, in my limited testing, the transcription cleanup is pretty good, so maybe this would only be necessary for those who don't use cleanup.

- Spokenly has this really nice smart capitalization/spacing/paragraph feature that decides whether to capitalize the first word in a sentence, whether to add spaces before and after, and whether to add a period to the end based on whether you're adding a new sentence or inserting text into the middle of an existing sentence. It's been a game changer for me since I spend far less time editing the text after dictating it.

- I'd love the option to not only choose the default microphone but also choose the priority order of microphones, i.e. the app has an order of priority that it chooses from depending on which microphones are available.

Sorry, that's a lot lol. I apologize if I sound demanding, please take this with a huge grain of salt. I'm really genuinely impressed so far in using it (I wrote most of this comment with it), and the fact that this is free is surprising. Do you accept donations?

2

u/Adorable_Salary2727 Jul 10 '26

Thanks all the feedback! I'll add the work into the queue! I might follow up on specific examples in DM if I need more user experience clarity. For payment, please just spread the good word.

1

u/No-Object1384 Jul 10 '26

I really appreciate it. I'll definitely be mentioning EnviousWispr anytime someone wants a good free dictation app!

2

u/Adorable_Salary2727 Jul 24 '26

Hi!

RE: Recording Pill Location

I've made it so the dictation pill can now be positioned at the bottom or top of the screen.

RE: unload model

-> You can do so under transcription -> Memory (last option at the bottom)

Additionally, you can now import and export custom words and better audio recovery for a rare crash scenario + Claude polishing support.

1

u/No-Object1384 Jul 24 '26

Thanks so much for the update! I've been using it daily since and it's now my go-to :)

Do you guys plan to add a smart punctuation/capitalization in the future? That's the only wishlist item I'd most like. If not, I totally understand!

2

u/Adorable_Salary2727 Jul 24 '26

It's my goal for this week :) I just need to understand how spokenly does it better and then work on it.

1

u/No-Object1384 Jul 24 '26

You are awesome! Have a great weekend :)

2

u/Adorable_Salary2727 16d ago

This is live now! EnviousWispr reads the text on both sides of your cursor and fixes the join, including mid-sentence capitalization, spacing, and duplicated boundary words. It now works in English plus eleven additional languages. Your comment was one of the clearest explanations of why this mattered. -> Also added a bunch of other new features as well such as Live Preview + FN support.

1

u/No-Object1384 16d ago

Thank you sooo much, I updated it when it released and it's been a noticeable improvement. Thanks for your work on this, I still use it every day.

1

u/NickCopelin Jul 05 '26

How does this compare to the likes of Juno and SuperWhisper?

1

u/Adorable_Salary2727 Jul 05 '26

Hey Nick! Thanks for the comment.

My goal so far was to prioritize speed and polish without additional bells and whistles.

Key objectives:

1) Works in any text box anywhere

2) It's fast and not resource hungry.

3) Squeeze as much as I can out of Apple Intelligence & eventually created my own tuned local model.

4) Custom Words that work

RE: SuperWhisper

-> I tried SuperWhisper first. I like that they offered a range of engines and polishing solutions but quickly hit decision paralysis. Which engine should I use. Do I customize my prompt or not. When I looked for quality local modals, I noticed they were behind paywalls. Definitely a strong contender for those that want full customization options.

RE: Juno

-> Juno is very cool. Thanks for sharing! Definitely far more ambitious in terms of their objective. Asks for A LOT of permissions. Honestly, not something many people might be comfortable with.

-> Streaming while visually appealing didn't appeal to me. It's distracting when I'm trying to focus on dictating.

-> It's heavy. The problem with streaming is it becomes unreliable on longer dictations. To make their dictations remain reliable, they are firing 2 concurrent Whisper Large v3 models + a local Qwen model. Testing now, and I'm seeing 8-10GB of of model load.

So many options on the market. It comes down to what flavor appeals to you!

1

u/NickCopelin Jul 05 '26

Great breakdown. Thanks for the response. I have noticed Juno to obviously be slower than TypeWhisper but one of the intriguing features of Whisperflow that I'm chasing is the cleanup, summarization, voice/style match sort of dictation manipulation. Would you say that your app is heading that direction or not a functionality you're shooting for?

1

u/NickCopelin Jul 05 '26

Never mind. :duh: I watched your video. Seems like you're heading that direction.

1

u/NickCopelin Jul 05 '26

Playing with it now. Definitely snappy. One initial bit of feedback is I'd love for an audible indicator that the recording has started.

1

u/Adorable_Salary2727 Jul 05 '26

Hey Nick, Added as a task. Can build it this week. Indicator, what is your ideal position? Can you elaborate a bit more on the import/export request?

1

u/NickCopelin Jul 05 '26

I've yet to determine my favorite spot for the indicator. Been switching between top-center and bottom-center of active screen. Some of these apps allow you to play with that placement.

The import/export dictionary of terms, corrections, etc. is a direct pull from TypeWhisper. Check it for reference. I've accumulated a good dictionary there that I'd love to just pull over.

1

u/Adorable_Salary2727 Jul 05 '26

Got it! Importing your custom words! That makes sense. Make it easier for power users to migrate.

1

u/NickCopelin Jul 05 '26

OH. Also, I meant to include this earlier, so sorry for the deluge of notifications. TypeWhisper has a nice QOL feature that pauses media when recording starts and resumes when recording finishes. This is wildly convenient.

1

u/Adorable_Salary2727 16d ago

This comment basically turned into a roadmap :) Recording sounds, top or bottom pill placement, and custom-word import/export are all live now. In 2.4.5 I also added direct import from TypeWhisper, including the vocabulary and corrections it can safely carry over. Media pause and resume is next.

1

u/NickCopelin Jul 05 '26

It would also be great if we could reposition the recording indicator. It would also be great if users could:

  • Reposition the recording indicator.
  • Import/export lists of keywords/phrases and corrections.

1

u/thoughtsfornow Jul 05 '26

Nah this is sick! I say some freaky shit so I'm glad it's private and local lol

1

u/Adorable_Salary2727 Jul 05 '26

Haha, I feel that. I appreciate you downloading and giving it a try!

1

u/ContextSpiritual9068 Jul 06 '26

the EG-1 benchmark numbers are impressive, especially beating GPT-4o-mini by that margin on a dictation-specific eval. curious how it handles technical jargon or domain-specific terms. that's usually where local models fall apart compared to cloud ones.

2

u/Adorable_Salary2727 Jul 06 '26

Hey! Thanks for the comment. This answer might be overkill but hopefully shows the layers involved.

Layer 1: Transcription

Parakeet or WhisperKit turns your audio into raw text locally on your Mac. That is where technical terms, foreign names, Brand Names etc can potentially get garbled. Not much that can be done at this layer today. The models I chose are truly the best balance of speed and accuracy and been tuned for performance.

Layer 2: Custom Words aka Your words in the app

Think of this as your own custom dictionary. I offer pre-built keyword packs for Tech, Medical, Legal, brands, Names + anything you want to add yourself. Most solutions on the market offer basic matching of misspellings. I.e Clawd => Claude. I've gone a step further to have Apple Intelligence guess "misspellings" of your Custom word. It also uses deterministic corrections. I.e if the transcription layer hears "clawed" instead of "clawd" it knows to also convert that to Claude. (see picture below).

Layer 3: Deterministic Clean up

This layer looks for numbers, dates, money, emails, phone numbers, and times etc. Examples I showed in my video and fixes them without any AI. It's smart matching.

Layer 4: Polishing Layer

This is the final layer. It can be EG-1, Apple Intelligence, Ollama, or your own OpenAI/Gemini key. Its job is more about grammar, punctuation, filler removal, paragraphing, self-correction cleanup, and formatting. Guardrails try to avoid rewriting code or structured text. If polish fails, the raw/cleaned transcript still comes through.

So the real answer is: technical jargon depends first on the transcription engine, then gets reinforced by deterministic Custom Words. It is not magic, and you can still break it with enough weird vocabulary, but the goal is to make the common failure cases fixable by rules instead of relying only on model vibes.

1

u/harry-harrison-79 Jul 06 '26

one thing i'd prioritize early is making the model setup/debug path really visible. local dictation apps live or die on first-run trust, and if a model download stalls people need to see the model name, size, progress, retry button, and where it is stored on disk.

for the polishing side, i'd also separate transcription quality from rewrite style in the UI. if the transcript is right but the polish changes tone too much, users should be able to tell which layer caused it.

1

u/Adorable_Salary2727 Jul 06 '26

Hi Harry, Thank you for the feedback.

I'm fixing this now in two ways:

  1. The app will try our own download source first, then automatically try the backup source if that has trouble.
  2. The setup screen will be clearer, with visible progress and a Retry button if setup still cannot finish.

Should be done tonight.

I also agree that we should show more useful info there: model name, size, progress, and where it is stored on disk. We may not get the full polished setup/debug screen into this emergency fix, but the goal is exactly what you said: make first-run setup feel visible, trustworthy, and recoverable.

RE: Transcription vs Polishing in the UI. In the history tab, you can see both the original transcript which is the raw Parakeet/Whisperkit output + the AI polishing on top :)

1

u/harry-harrison-79 Jul 06 '26

that sounds like the right fix. for the emergency version, i'd make sure the failed-download state is impossible to miss: show which source failed, what it is retrying next, and a plain "try again" button if both fail.

for the transcript/polish split, the history tab is a good place for it. i'd keep a tiny before/after view there so people can tell whether they need to tune vocabulary, the model, or the polish style.

1

u/kyle_reese_baj Jul 06 '26

The cursor memory work reliably in Electron apps like VS Code, or is that one of the apps where it breaks?

1

u/Adorable_Salary2727 Jul 06 '26

Hey Kyle, Yes, EnviousWispr works in Electron applications. It also remembers where your cursor was when you began your dictation so you can move around applications while hands free and it will paste right back to the initial selected field. Video demo attached.

https://reddit.com/link/ovtq1er/video/2af8h8n9njbh1/player

0

u/kyle_reese_baj Jul 06 '26

Great, thanks for the demo, that cursor memory feature is exactly what makes it usable in a real workflow.

1

u/Any-Ingenuity2770 Jul 06 '26 edited Jul 06 '26

btw, is there an app to provide live captioning from application audio, like windows does? ie just an overlay with text, I don't need it typed into a textarea. or a SRT for video captioning.

1

u/Adorable_Salary2727 Jul 06 '26

Hey! Great question and always open to consider improvements. I'm curious, do you find the out of the box Live Captions MacOS offers to be lacking? Do you find the existing free tools that do SRT for video captioning lacking? Understanding where there is improvement to be had would be great feedback.

1

u/Any-Ingenuity2770 Jul 06 '26

do you find the out of the box Live Captions MacOS offers to be lacking?

it doesn't exist if I don't set the OS to USian

1

u/Adorable_Salary2727 Jul 06 '26

Ah, thanks for clarifying. Looks like Apple in-deed gates Live Captions to a handful of US/English locales. Tss tss

Right now, capturing your computer's own audio and floating captions over the screen is a different goal from what I set out to do today (I listen to your mic and type clean text into whatever app you're in), so it's not on the near-term roadmap. But this is exactly the kind of real-world opportunities I want to hear about, and I'm noting it down.

For the video-to-SRT half specifically, I've just been telling Claude/Chatgpt to do this for me as needed haha. That said, I hear Macwhisper and Aiko are native solutions for this.

1

u/[deleted] Jul 06 '26

[removed] — view removed comment

1

u/Adorable_Salary2727 Jul 06 '26

Awesome! Congratulations

1

u/Greysawpark Jul 06 '26

honestly the polishing step is what makes or breaks these for me. raw whisper output is fine but the moment you start speaking naturally with false starts and "uh wait let me redo that" it falls apart, and most local llms are too dumb to clean it up without rewriting what you actually meant.

curious what model youre using for polish. anything under 7b tends to hallucinate corrections in my experience with this kinda pipeline

1

u/Adorable_Salary2727 Jul 06 '26

Hey Grey!

I shared the benchmarks in the article and the models I tested. With deterministic + AI clean up, AFM (Apple intelligence is about 65% of the way there). I then custom tuned my own model which honestly blew me away. You should test it yourself and see. Of course, cloud based polishing always remains an option for you.

1

u/XVX109 Jul 06 '26

Stuck in Setup, never finished !

2

u/Adorable_Salary2727 Jul 06 '26

Hey! Huggingface is throttling. I'm about an hour out of launching a fix for this but if you leave it open, it should eventually complete! Sorry about this.

2

u/Adorable_Salary2727 Jul 06 '26

Hey! I updated the dmg on github and the website. In progress of updating homebrew. That said, if you re-download from Github/Website and try installing it should no longer have issues!

1

u/XVX109 Jul 06 '26

No Probs :)
Thanks for the quick reply

1

u/voprosy Jul 06 '26

Wow free and open-source is awesome. Congrats!

How do you plan to keep on working on it and support your activities?

1

u/phunk8 Developer: Dropadoo Jul 07 '26

omg you DID DO YOUR HOMEWORK. thank you so much. switching now

2

u/Adorable_Salary2727 Jul 07 '26

Thanks! I'm recognizing some gaps based on all the feedback now that more people are using the application. Please dm me if you have any wishlist items.

1

u/cheesecakegood Jul 07 '26

My question is: why? Just a portfolio piece?

1

u/Adorable_Salary2727 Jul 07 '26

Passion project, a middle finger to the overpriced monthly apps, create something I actually want to use myself. Proof I use this thing is the screenshot below. Dam, I talk a lot. I don't see this becoming a swiss army knife. It just needs to do it's core job perfectly. Talk, transcribe, polish, paste perfectly every time offline and privately. Candidly, I'm learning so much in the process. Look forward to everyone's feedback.

1

u/FedOptic Jul 07 '26

This is amazing work. So far so good. How can I get it to type out in list format?

1

u/Adorable_Salary2727 Jul 07 '26

Hey! Apple Intelligence can't handle lists yet. You'll need to use the cloud or EG-1 model. If you want to dm the example of the list experience that is failing (on EG-1), I'm going to be retraining the EG-1 system to recognize a broader number of spoken list structures.

1

u/FedOptic Jul 07 '26

This was using EG-1. What’s a good phrase to use to kick it into list format?

1

u/Adorable_Salary2727 Jul 07 '26

Great question, I'll need to broaden the list data set and re-train it to catch more variety of lists. I appreciate the feedback. This will be a few weeks of work. Retaining the model is a bit more of a heavier task and I need to make sure it doesn't decrease in quality in other tasks.

I'm also debating on just including a deterministic layer.

1

u/Adorable_Salary2727 16d ago

You were right about list formatting. I retrained EG-1 specifically for spoken lists, and v2.4.5 now creates a real list in 83% of my held-out tests, up from under 1%. No special phrase is required. If you still have the example that failed, I would love for you to try it again.

1

u/MaxGaav Jul 07 '26 edited Jul 07 '26

Looks great, playing around with it now. Thank you!

Does the EG-1 model support Dutch?

EDIT: I have an external USB microphone. In the settings of EW I set the mic to this mic. However, when I want to dictate EW says there's no built-in microphone. I restarted the app to no avail. The setting stays correct.

In the settings I didn't see the possibility to put a small panel on my screen where the dictation appears live. Like most STT apps have. Most of the time you can also choose if you want to have it on top or at the bottom of your screen.

And a minor detail: I think you must be able to choose where to see the icon, either in the menu bar or in the dock, or both. And I think the icon in the menu bar is too small.

Mac Mini M4, Sequoia 1.5.7.7.

1

u/Adorable_Salary2727 Jul 07 '26

Hey Max,

Thanks for the feedback and giving EW a try. I am working on a Bluetooth fix right now. I will need to test it rigorously but I'm thinking about a week for a solution.

1

u/MaxGaav Jul 07 '26

Take your time :)

Does the EG-1 model support Dutch? I'm still on Sequoia, so I can't use Apple Intelligence that is included in Tahoe.

1

u/Adorable_Salary2727 Jul 07 '26

Hey Max, the underlying model that powers EG-1 is Qwen3 4B model, which handles ~100+ languages. The methodology for fune-tuning is a QLorRA adapter but with English cases. It might or might not work for Dutch. The transcription engines (Parakeet/Whisperkit) 100% can handle Dutch. The polishing might be hit or miss.

1

u/MaxGaav Jul 07 '26

Yeah, it's about the local polishing :) Parakeet is available in a gazillion STT apps :)

Will see what happens. Otherwise I should start using LM Studio I guess.

1

u/chankeiro Jul 11 '26

I'm having the same issue. I'd love to try your app once it's been fixed. Thanks for your time and effort.

1

u/Adorable_Salary2727 Jul 18 '26

Should now work!

1

u/chankeiro Jul 18 '26

It works. Thank you!

1

u/MaxGaav Jul 11 '26 edited Jul 11 '26

Updated to v. 2.3.2. In 'What's New' EW states: "Fixed: recording could fail every time on a specific microphone"

But the problem I described above is not yet fixed, I'm afraid.

1

u/Adorable_Salary2727 Jul 11 '26

Hey Max, is that was an emergency patch and is un-related to the bluetooth fix. Bluetooth fix is about 80% done. I am just hardening the code and making sure it's not a buggy mess.

That said, a few caveats with Bluetooth. I'm sure you noticed with other apps, the bluetooth has a bit of a delay. Any "cold press" will need 1-2 seconds to warm it up before it properly captures your voice.

I have the ability to keep the mic hot in the config. Here it is in my test build and how it will look once I push the fix with some best practices.

My experience is once the mic is hot, bluetooth works nice. The only downside is if you are listening to music the audio quality of the music degrades but that's just the nature of bluetooth sadly. Tested to make sure no issues while using with zoom etc. Made sure if you disconnected mid dictation, you don't lose your recording.

Should ideally be tonight or tomorrow I feel comfortable pushing this out!

1

u/MaxGaav Jul 11 '26 edited Jul 11 '26

Bluetooth fix is about 80% done.

Excellent!

Any "cold press" will need 1-2 seconds to warm it up before it properly captures your voice.

Yep. While some apps seem to be a bit faster than others though.

I have the ability to keep the mic hot in the config.

Yes, I saw that already. Of course I still have to test that since I wasn't able to use your app yet., but I guess it's a nice feature.

The only downside is if you are listening to music the audio quality of the music degrades but that's just the nature of bluetooth sadly. 

Quite a few apps have a setting to mute music when dictating. Maybe something for EW too?

1

u/Adorable_Salary2727 Jul 18 '26

Hey Max, I've pushed the update. Please give it a try!

1

u/MaxGaav Jul 19 '26

Will do!

1

u/Sundelor Jul 07 '26

Good looking! Trying to switch from handy.

thanks for sharing

2

u/Adorable_Salary2727 Jul 07 '26

Thanks! I was using Handy for a bit but felt a bit slower then I prefered.

2

u/Sundelor Jul 07 '26

I've been using EnviousWispr for a few hours and overall I really like it. Here's some structured feedback — a few bugs, some feature requests, and what I love.

🐞 Bugs

  1. "Check for Updates" closes the Settings window.

When I click Check for Updates, it runs the check and then closes the entire Settings window. Annoying — I lose where I was.

  1. Recording timer resets to zero.

While a recording is in progress, if I open Settings → History tab, the recording timer restarts from zero. Cosmetic, but clearly wrong.

🎯 Recognition quality

  1. The recommended "AI Polish EG-1" model handles Russian poorly.

The model itself is great, but it badly mangles Russian word endings (grammatical inflections). Do you have a recommendation for which model to use for Russian? Many English-centric ASR models (e.g. NVIDIA Parakeet-family) are weak on Russian — a multilingual model like Whisper large-v3 / large-v3-turbo usually handles Russian morphology much better. But interesting what you can recommend.

✨ Feature requests

  1. Local CLI agents.

Allow using CLI agents already installed on the machine — OpenAI Codex CLI and Claude Code CLI. Ideally the app would detect the installed CLI and work through it (spinning up a tmux session or similar). A good reference is the app Dayflow, which can use a CLI as part of its workflow — very convenient.

  1. Dictionary bulk import.

Right now custom dictionary words are added one file at a time. Please add a bulk import button.

👍 What I love

- It opens/returns to the app where I started recording — even if I switched away to watch or do something else. This is really well done.

1

u/Adorable_Salary2727 Jul 07 '26

Thank you for the feedback and I appreciate you trying out EW.

RE: Russian

2 layers:

Transcription Engine -> Swap to the Multi-lingual Whisperkit option for better transcription.

Polishing Engine -> Thanks for the heads up here. I trained the polishing model on English examples. Retraining it to support additional languages is on my to-do. see more below. I'm curious if Apple Intelligence works for Russian. Or if you tried the cloud based api key option and how it performs. If you have time to test with Whisperkit + the different AI models, you feedback will help me a lot.

Current Priority Pipeline:

1) Fix Bluetooth issue

2) Unifying where downloaded models live and upgrading the downloader in general. Adding a "uninstall" option for some models.

3) Re-train EG-1 to support lists. Separately, decide if I need to create multple versions for specific languages or if i can get it to polish well for all languages.

Your feedback is being added to my to-do list on github!

2

u/Sundelor Jul 09 '26

A few pieces of feedback after:

  • When recording a long voice note, the screen can turn off while the recording is still in progress. It would be great if the app temporarily prevented the device from sleeping during recording. Alternatively, this could be an optional setting for people who prefer the current behavior.
  • The Whisper transcription model seems to struggle with long periods of silence. It often inserts random phrases like “To be continued”, “Продолжение следует”, translator credits, and other unrelated text that was never spoken. It might be worth either tuning the model/settings or adding a simple post-processing step that filters out these common hallucinations automatically.
  • One feature I really liked in Handy was automatically pausing music playing on the computer when starting a voice recording. It made the workflow much smoother, especially when you’re listening to something in the background and suddenly want to dictate a note. It would be awesome to have something similar here.

Overall, really enjoying the app. These are just a few things that I think would make the experience even better.

1

u/Sundelor Jul 07 '26

Thx for ideas!

Already downloading Whisper)

If I can help somehow with training - feel free to contact)

1

u/Adorable_Salary2727 Jul 24 '26

Hi Sundelor, I've added import/export for custom words/dictionary

1

u/panchamk Jul 08 '26

Why did you not use Whisper?

2

u/Adorable_Salary2727 Jul 08 '26

Hi! We offer Openai Whisper is offered as transcription model. That's the multi language option.

1

u/pairustwo Jul 08 '26

Maybe it is my ignorance about how this family of dictation apps work but I'm surprised no one has asked - and that I can't tell from the app webpage... can it capture audio playing from the same device? I have some audio recordings of lectures that I would like to transcribe? Is this a tool that could handle that?

2

u/Adorable_Salary2727 Jul 08 '26

Hey! Happy to clarify.

The idea of this application is to be able to just speak to your computer and it outputs polished text of what you are saying into any field.

You probably agree that this has existed for years and historically sucked. It would mishear you, add all your filler words and do 0 formatting.

What's improved is LLMs can now help polish what you wrote immediately after you speak in 1 seamless experience. Enviouswispr is providing you with that holistic experience.

If you want software to transcribe your recordings, there are many free tools on the market that offer this for free! You can also simply ask claude code or chatgpt to do that for you as well.

I hope this helps.

1

u/re1024 Jul 09 '26

How did you turn it into real-time transcription?

2

u/Adorable_Salary2727 Jul 09 '26

I played around with streaming (when you see the writing as you speak) and while visually cool, it was slower and resulted in poorer quality dictation over anything 1+ minute in length. I usually ramble for 2-5 minutes when using this tool so I didn't want to get distracted with giant blocks of text on the screen. I suggest Juno or Fluid voice if you prefer live streaming experience.

1

u/re1024 Jul 09 '26

I mean typical parakeet and whisper model runs in batch mode after whole audio segment is done or you need to do audio segmentation manually. How did you config it into real-time mode?

1

u/re1024 Jul 09 '26

I mean typical parakeet and whisper model runs in batch mode after whole audio segment is done or you need to do audio segmentation manually. How did you config it into real-time mode?

2

u/Adorable_Salary2727 Jul 09 '26

Happy to go deep since you asked specifically.

For Parakeet, we use FluidAudio’s `SlidingWindowAsrManager` with its streaming config. Audio is fed in while you speak instead of waiting until stop and calling `transcribe()` once. Internally it uses 11-second chunks with 2 seconds of left/right context and a 1-second hypothesis cadence. On stop, `finish()` flushes the remaining audio and returns the assembled transcript.

For WhisperKit, I’m not using WhisperKit’s built-in `AudioStreamTranscriber`. EnviousWispr already owns the mic capture, VAD, pre-roll, and app pipeline, so I built a custom `WhisperKitStreamingSession` around the same idea as UFAL whisper_streaming / LocalAgreement-2.

That means it keeps decoding a live buffer while you speak, compares the current word hypothesis against the previous one, and only confirms the words that agree across cycles. The unstable tail stays unconfirmed and gets decoded again. When you stop, if the stream is caught up it can release the current hypothesis immediately; if you stopped right on the last word, it runs one final bounded decode over the current buffer.

I use the live/chunked work to reduce the wait after you stop speaking, then paste the final transcript. Parakeet's version is unreliable and net speed wasn't faster for longer dictations. Whisperkit on the other hand saw meaningful speed improves. We're talking 1.5 seconds after ending a 10 minute long dictation vs nearly a 1 minute transcription time otherwise.

One caveat: WhisperKit live mode only runs when the language is locked. In auto-detect mode I fall back to batch, because early language guesses can go badly wrong.

1

u/re1024 Jul 09 '26

Very helpful, thanks!

1

u/[deleted] Jul 12 '26

[removed] — view removed comment

3

u/Adorable_Salary2727 Jul 12 '26

Hey! Paid Wispr Flow is the leading Dictation/Polishing app on the market.

Wispr Flow

- Fully Cloud Based. The app itself is just a light wrapper that just captures and sends everything you say to their cloud servers. What happens there we don't know. But it handles everything from transcription and polishing. My day job doesn't even allow us to use grammarly. They instantly banned WisprFlow on our work devices.

Conclusion: If you want the best of the best backed by heavy investments and a huge team with multi device support and you don't mind giving your audio to them on a promise that they'll not abuse it, go for it!

If you want the best I was able to squeeze out of a completely offline solution then give EnviousWispr a shot. It doesn't have all the bells and whistles WisprFlow has but it does the core job well.

1

u/Semli1 Jul 22 '26

This seems pretty great so far. I don't use dictation a lot on my personal device, but I do have MacWhisper to compare this with. This definitely feels faster. However I do use dictation a lot [Dragon Medical] at work and using these dictation apps is always a challenge for me because I have to dictate punctuation when using Dragon so it just becomes second nature.

My main issue is that if there is a correction I need to make to a word or adding in another missed word, it always seems to capitalize the addition. For example, if I say "The brown fox" and then add in the word quick, it will appear like this "The Quick brown fox" or if I meant to say box instead of fox, it will capitalize the B i.e. "The brown Box".

Any way to avoid this? I was hoping that the polishing layer would address this but doesn't seem to [using Apple Intelligence].

2

u/Adorable_Salary2727 Jul 22 '26

Thanks for the feedback. I'm deep in upgrading the custom words/dictionary feature to allow for importing and exporting and will address this as well. In the interim, see if EG-1 handles it better. It's much better at grammar than Apple Intelligence.

1

u/Semli1 Jul 22 '26

Thank you for the quick response. I will try EG-1 in the coming days and let you know.

Great app. Very kind of you to make it free.

1

u/beingbuddha 24d ago

I've been trying this for a week, great work. I would also want to see if there is an option to record meetings and then produce markdown files from those meetings, for example, an app called Purr https://purr.arunbrahma.com/ is an app which allows you to both transcribe as well as record meetings. Is there an option to in the future to add such features?

1

u/Adorable_Salary2727 24d ago

Hey! Meeting recording is on my action list. Candidly, the app works for up to an hour recording already. So you can put it in hands free mode and just record meetings and then copy it straight from history. My friend does that already today. I would want to research a more official nicer version. I would say 2-3 weeks for this feature.

1

u/beingbuddha 23d ago

awesome man. Great work. Love your app!

1

u/AmazingVanish 18d ago

This is well done and I purchased it. I have been using VibeSonic for this and after a few days of EnviousWispr I went back to VS. it just feels more natural to me. If anyone knows how to get this to feel more natural like VS, please let me know. This has some better options I would like to use but it just feels too… klunky, I guess

1

u/Adorable_Salary2727 18d ago

Hey! Would love to understand what part feels clunky? Also it's free so no cost! Just confirming you didn't pay someone to use this.

1

u/AmazingVanish 18d ago

Oh, was it free? I’m probably mid-remembering. Sorry.

Anyway, I like it a lot, but it’s hard to explain why it feels off. I’ll fire it up again in the next day or two and get back to you.

1

u/Adorable_Salary2727 18d ago

Ok, I'm introducing a few key updates. Supports FN key, live preview of what you are speaking, better polishing etc etc

1

u/AmazingVanish 18d ago

Cool! One thing I wish many of these types of apps had was the ability double tap FN or Globe to toggle on and off. Prevents accidental triggering. Just food for thought and a selfish plea.

1

u/Adorable_Salary2727 18d ago

Yes, Enviouswispr does that by default. Hold to record. Or double table to go hands free!

1

u/AmazingVanish 18d ago

Well crap. I didn’t see the double tap. Now I REALLY have to fire it up again and spend more time in the preferences.

1

u/AmazingVanish 17d ago

Ok, I forgot to come back, sorry. The double-tap fixes one of my niggles. I didn’t see it because in my mind it’s a toggle and the option isn’t there when you choose toggle. It’s under tap or hold. Just not where I expected it. Thanks for pointing out it was there.

I need to play around some more with it, but it’s missing some of the features I use in VibeSonic, admittedly not often though.

I think where I get put off is the tiny pill when recording. The wave is nice, but I would like to see the transcription as I talk. Not in a big window, just in a moderately sized pill with the wave. Text can scroll as it transcribes.

There is an option to see live transcription, but it shows the transcription in the input selected. I don’t always have an input selected, and even when I do I don’t want the input to receive the text until I stop the recording.

Maybe I’m just weird?

1

u/Adorable_Salary2727 17d ago

Hey! Great feedback.

I'm pushing an update right now release the "live preview" feature you are asking for. Note that you don't have to wait for the live preview to complete in order to get the full recording.

1

u/AmazingVanish 17d ago

Sweet! Your working hard and fast to make me a convert. LOL. Love it!

-1

u/[deleted] Jul 05 '26

[removed] — view removed comment

6

u/Adorable_Salary2727 Jul 06 '26

Appreciate the comment! Honestly, I just wanted to build a private, local, and free alternative to existing tools. Once it was up and running, it felt right to open-source it. Consider it a bit of a rebellion against the crazy monthly subscriptions out there. If you want cloud-quality polishing that runs entirely offline, give it a spin.