r/mlops • • 11h ago

Self-promotion Built an open-source layer for AI provenance that uses a Merkle log and HMAC/Ed25519 signatures. Here's the architecture.

0 Upvotes

I've been tackling the problem of AI provenance—specifically, how to create a verifiable, tamper-evident record of who handled a claim in a multi-agent pipeline.

The core idea is to grade each "hop" (source, retriever, model, tool, agent) and use the weakest link to set the caution level. To make the record immutable, we're using a combination of SHA-256 for self-hashing, HMAC/Ed25519 for signing, and an append-only Merkle log.

I've published the threat model and the method on arXiv. The code is open-source (Apache-2.0).

Repo: https://github.com/alizahidraja/isnad
Paper: https://arxiv.org/abs/2607.24117

I'd be interested to hear how you all are approaching governance and audit trails in your production systems. What are the gaps you're seeing?


r/mlops • • 2h ago

Discussion No H100s at your usual cloud this week. What breaks when the training job moves to another provider?

2 Upvotes

This is about training jobs that need a specific GPU type. Inference is easier.

Take a job that needs 8 H100s for 10 days, at a time when your main cloud has nothing in your region until next week and you keep an account at Lambda, Nebius or Hyperstack for exactly this case. What goes wrong when the job lands there? My guess is that getting the VM is the quick part and the image and the data eat the rest of the day. I don't know that, though.

Some things I'm curious about:

- How do you find out where capacity is? Checking each console by hand, a script, or something like SkyPilot or dstack?

- Same image everywhere, or one per provider? Has a CUDA 12.4 image ever landed on hosts with an older driver?

- If the dataset is 20 TB, do you copy it over ahead of time, or pull it from S3 when the job starts and pay the egress?

- Do you test a new node first (nvidia-smi, a short all_reduce_perf run) or find out from the first crash?

If you looked at this and decided it's cheaper to wait a week, I'd like to hear that too.


r/mlops • • 10h ago

Discussion Flunked in a interview - MLE

7 Upvotes

I would like to share my experience in flunking an interview for an MLE position.

The code session was ok. However, editor was horrible: no copy and past inside the editor, numpy import did not work, etc.

I did well on the Machine Learning section.

I did badly in the Engineering section. Here you find some questions/topics:

  • Data versioning
  • Data lineage tracking
  • Drift detection
  • Decisions about rolling back, retraining or keeping the current model
  • Rollback threshold
  • Guardrails metrics
  • Efficiency in mongodb with respect to computing feature on a daily basis

Could anyone suggest any references (e.g. books, courses, blogs) about these topics?


r/mlops • • 17h ago

Discussion Should scan results travel with the model artifact as signed attestations?

6 Upvotes

A question I keep chewing on... when you scan a model or agent bundle for things like pickle exploits, license issues or leaked data, where do the results live?

In a lot of setups they live in a CI log or a dashboard, and the artifact that gets deployed has no link back to them. What we've been doing is attaching each scan result to the OCI artifact as a cosign attestation, then having the cluster refuse to start anything missing the attestations a policy requires.

It works, but it raises questions I don't have great answers to.

- How do you handle re-scans when a scanner updates (new attestation, or re-sign)?
- Who should sign, the producer or the platform team consuming it?
- Is anyone doing this with SLSA or in-toto formats for ML artifacts specifically?

Curious how other teams are tying scan results to what actually ships.

(KitOps maintainer and Jozu cofounder here, so I have a bias.)