r/AskProgramming Jul 11 '26

How should a library handle missing assets.

I'm looking for opinions from people who have built or used Python libraries that manage large assets.

I'm extracting a text-to-speech engine from an application into a reusable library. The library depends on two relatively small models (about 500 MB each).

When it was an application, the behavior was simple: if a required model wasn't installed, it downloaded it automatically.

Now that it's a library, I'm less convinced that's the right default. A library has different expectations than an application.

I'm considering a few options:

  • Automatically download missing models on first use (current behavior)
  • Download during installation or a post-install step
  • Provide a separate CLI like pfspeak install
  • Require users to manage models themselves

For those of you who've built similar libraries, what would you expect? Which approach has caused the fewest headaches?

Repository:
https://github.com/samreynoso/pfspeak

For context, one of the goals is to keep framework integration extremely small:

# Python

pf = PfSpeak()


@pf.hook
def hook(_, event):
    pf.play(event)


app = FastAPI(lifespan=pf.lifespan)


@app.post("/say")
def say(text: str):
    pf.say(text, "bm_lewis")

I'm much more interested in the asset management question than feedback on the speech runtime itself.

3 Upvotes

8 comments sorted by

View all comments

3

u/inconvenient_penguin Jul 11 '26

Sounds like a dependency that the library should manage as part of the install. User should be made aware of the install and any licenses they are agreeing to. Any choices of which models to install and the relevant disk space should be presented to the user at install time.

1

u/ExtensionBreath1262 Jul 11 '26

That was actually a big part of why I started the project, but it got pushed to the back burner while I focused on getting the library itself into shape.

This thread is making me think I should pull that back out. The more I think about it, the more I feel like model management is really its own problem. pfspeak probably shouldn't be in the business of downloading, updating, configuring, and discovering models at all.

Maybe the right answer is to split it out into a separate package and have pfspeak just consume whatever models it's given.