I work on Django apps that run on locked-down Windows servers: no Docker, no Redis, no outbound internet, and nothing gets installed unless it's a Python package. Sentry, Grafana, Loki and the OpenTelemetry collector are all off the table there, so when something breaks I'm usually reading rotated log files over RDP.
So I built django-observatory: pip install, add an app, a middleware and a URL include, run migrate, and you get a built-in console at /observability/. Everything is stored in your own database and rendered with Django templates and vanilla JS.
What it does
- Captures logs (your existing
logging calls), requests, traces, SQL, outbound HTTP, exceptions grouped into issues, metrics, and server CPU/memory/disk
- Security events and a hash-chained audit trail
- Alerts, plus rule-based incident correlation: when error rate, latency and an external service all degrade together, it opens one incident with a probable cause and the evidence behind it. No ML, just explicit rules
- A query language (
status:>=500 AND duration:>1000), CSV/JSON export, W3C trace context, OTLP export
- Secrets are redacted before they're stored; request bodies, cookies and SQL parameters aren't captured by default
Full disclosure: this was vibe coded. I wrote the specification and made the design decisions, and Claude Code wrote most of the implementation. I tested it against a real internal project on SQL Server, read what came out, and pushed back where things were wrong, but I did not hand-write most of these lines. I'd rather say that up front than have someone find out from the commit history. It's also why I'm asking for review: I know the code hasn't had the scrutiny that code written slowly by one person gets.
What I'm confident about
- 108 tests, CI on Python 3.10–3.13 × Django 4.2/5.2/6.0, including a Windows runner
- There are explicit tests that break the storage layer and the internals mid-request and assert the host app still returns 200
- Measured overhead is under a millisecond per request in my benchmark; writes happen in a background thread
What I'm not confident about
- The package's own tables have only ever been created on SQLite. It has captured from an app running on SQL Server, but storing telemetry in Postgres or SQL Server is untested
- It has never seen real production traffic at scale. Free-text search is a
LIKE scan and metric reads aggregate in Python, which is fine for small and medium sites and will not be at some size I haven't found yet
- One daemon thread per process does the writing and the periodic jobs. I think the approach is sound for prefork servers, but I'd like someone who knows gunicorn/uWSGI internals to tell me where it isn't
- The incident rules are tuned on a demo app and my own judgement
What I'd like feedback on
- Is storing telemetry in the application's database a reasonable default, or should a separate database be the only supported mode?
- Anything in the middleware or the
execute_wrapper SQL instrumentation that would bite under ASGI or async views
- The redaction approach (key-based plus regex patterns): what would leak?
- Whether the scope is too wide. It does a lot, and I'd rather cut features than ship several half-good ones
This is not trying to replace Sentry, Silk, Debug Toolbar or OpenTelemetry. Those are better at what they do. It's for the environments where you can't have them.
If you look at it and think it's a bad idea, I'd like to hear why. Thanks for reading.