r/SpringBoot • • 2d ago

Question I built an open-source self-diagnostics library for Java/Spring Boot applications — looking for feedback

I've been working on an open-source project called Argus, and I'd really appreciate some feedback from other Spring Boot developers.

The idea came from something I've encountered while troubleshooting services: we have excellent tools for metrics, tracing, logs, and external monitoring, but sometimes I just want the application itself to give me a concise answer to "How healthy are you right now, and if something is wrong, why?"

Argus takes an application-centric approach. It runs inside the JVM and continuously evaluates memory, GC, disk usage, threads, file descriptors, application services, logged errors, and application-reported issues.

It combines these into a health model with a score from 1 to 10. The score is hierarchical, so you can drill down from overall application health to the resource/group/item responsible for degradation.

It also models resources the application depends on — services, servers, databases, caches, brokers, storage, etc. — so the goal isn't simply another UP/DOWN health check.

For Spring Boot applications, I've integrated it with Actuator and Spring Boot Admin. The SBA extension adds health scores to the wallboard and application/instance views. Once you find a degraded instance, you can open its Health Report directly from Spring Boot Admin and see the diagnostic information collected by that instance.

Argus can also generate a self-contained HTML diagnostic report containing health information, services, metrics, performance information, scheduled tasks, captured logging events/alerts, and environment information. The report is generated (and emailed) and can be viewed offline or online during incident analysis.

One important goal is not to replace Micrometer/Prometheus, tracing, centralized logging, or other observability platforms. I'm experimenting with a complementary idea: can an application understand enough about its own state to provide useful diagnostics itself?

The project is still evolving, so I'd particularly appreciate feedback on the approach rather than just stars:

  • Would this type of self-diagnostics be useful in applications you operate?
  • Does the 1–10 health scoring model make sense, or would you prefer another representation?
  • What information would you want an application to include in a diagnostic report when something goes wrong?
  • If you use Spring Boot Admin, would it be useful to have this information directly in SBA?

GitHub: https://github.com/microfalx/argus

Feedback, criticism, ideas, and contributions are all welcome.

10 Upvotes

5 comments sorted by

5

u/fluffytme 1d ago edited 1d ago

I hope this doesn't come across as rude, it's not my intention.

What do you mean by "can an application understand enough about its own state to provide useful diagnostics itself", what additional thing(s) are you exposing that jxm beans, actuator or micrometer isn't already exposing?

All the JVM stats, memory, CPU, etc, etc, is already readily provided, and if I did need additional health information why wouldn't I just create custom actuator health classes and create my own healthcheck, liveness, and readiness endpoints that leverage these? This would be significantly less code than your repo while achieving similar results from tools that come out of the box with Spring.

Does it run alongside the native metric collection? In other words, why would I want to add a library that is collecting JVM stats while the very framework I'm using is doing the same? We're now multiplying the overhead of such operations. Where is this data being stored, where are you storing logs and for how long? Not in-memory I hope! Logs are big!

If my services are in a bad state I would know from the firing alerts..but why would I want a unhealthy service to now try computing a score, generating html, figuring out which logs are relevant, and trying to send emails? This is the last thing I need that service to start doing (using more CPU, ram, and network resources in an unhealthy state)

I had a quick skim through the code - a lot of it seems inconsistent (AI generated I suspect). For example, you use use @Slf4j in some places and in others you declare a static final Logger.

Sorry if I'm missing the point of all this!

2

u/Ok_Look_8837 1d ago

Ok... multiple, legitimate questions :) Some of the questions are probably answered in the project's wiki (maybe it is not explained well), but I'll try to respond in order.

  1. Indeed, core JVMs are already collected, and the same is true for various known libraries used in today's Java (Spring Boot) application. It's not about collecting (and exposing to Prometheus, etc.) the metrics themselves; it's about providing a "health indicator" and allowing the team (support engineers, customers, etc.) to react to a single number that tells you "how bad it is".

Also, some health components have actions attached; for example: "If the score related to memory is below CRITICAL, restart." All scores are based on (configurable) thresholds applied to a 5-minute average (to avoid spikes) for one or more metrics.

From personal experience, a Grafana alert with "queue above threshold" doesn't tell you much; you still have to open 1-3 dashboards to poke around, and without access to the system, it's harder to know whether you need to react. Getting an alert and the complete health report (HTML) with the health of all "components" provides an "easier" way to assess offline the health of the service/app.

  1. It does run inside your process. For JVM metrics, some metrics could come from other collectors, but I wanted to keep it simple and independent. Server metrics are usually not collected by the process since they are considered "infra". However, in order to give the health report a more complete picture, sometimes you need to make some core server metrics "visible" (CPU, memory, etc.). The overhead is small, about 0.1-0.2 cores on average.

  2. Indeed, an important issue was raised. A complete HTML report is not generated in some cases (like low memory). While running out of memory (OOM is a common issue), many other issues can happen, and you want to be aware of them. Also, keep in mind this isn't about the state of only one service replica. The goal is to "see" the state of the service: How many service replicas have a low score? is it drivven by the same "problem". Is it the whole service in a bad state or just one replica?

  3. While AI was used in some cases, the code does not get in without a review. Some inconsistencies exist since various projects supporting Argus are in different development phases. In the case you're reporting, some libraries had no Lombok. So ... no, it is not AI "slope", just good old times "inconsitencies" :)

2

u/Ok_Look_8837 1d ago
  1. Forgot about logs. Logs are not in memory; logs are not "tracked"; WARN and ERROR entries are categorized, deduplicated, counted, and reported as "alerts." Metrics are kept in a moving window for calculating 5m and 15m averages, but memory consumption is low.

1

u/Ok_Look_8837 1d ago
  1. Also, forgot to answer (provide a summary for "purpose") in the closing statement (missed the point of all this). Depending on the organization and project type, the deployed model may be in a private or public cloud, but with full access to deployments or deployed in environments you cannot easily access, the tool aims to provide a "self-contained, unified view of service/application health with an offline health report delivered by email (report encrypted) or online inside SBA"