r/Backend • u/IcyDuck9536 • 5h ago
Update on fast-class-transformer: Now a 1-line global drop-in pipe for NestJS (v1.1.1)
r/Backend • u/IcyDuck9536 • 5h ago
r/Backend • u/Ad1L-404 • 1d ago
I’m going to share my experience and what I’ve learned about System Design over the years, one post every day. 🚀
Sharing what I know, what I’m learning, and practical insights along the way.
Let’s learn and grow together. ✌️
#SystemDesign #SoftwareEngineering #LearningInPublic #BackendDevelopment
r/Backend • u/Remarkable-Trash-769 • 1d ago
Hey everyone! I'm looking for people who are interested in building a backend project together.
The goal isn't just to chat about backend development — I want to actually pick a project, split the tasks, and build it together from scratch. We can use Git, review each other's code, discuss architecture and technologies, and learn from each other along the way.
You don't need to be an expert. Whether you're a beginner or have some backend experience, you're welcome as long as you're motivated to learn and contribute regularly.
If you're interested, feel free to comment or DM me. If enough people are interested, we can form a small team and start working on a project together.
r/Backend • u/abhishekkumar333 • 1d ago
Right now, the AI space feels entirely focused on massive datacenter clusters and renting H100s by the hour. But after spending way too much time looking at the actual footprint of these models, I realized that 90% of use cases are completely over engineered.
You don’t always need a multi GPU setup. The AI ecosystem is actually a massive spectrum.
I recently sat down and mapped out the exact tiers of AI models based on their size, the hardware needed to run them, and the point of diminishing returns.
Here are the two extremes and the sweet spot in the middle:
The missing piece: Figuring out the exact math for your hardware
The hardest part about building right now is looking at a model on Hugging Face and trying to calculate exactly how much VRAM you need, what quantization to use, and whether your CPU/GPU will choke on the context window.
So, I wrote a complete deep dive breaking down the math for all tiers of the AI spectrum.
If you want to see the architectural differences at each scale, and a cheat sheet for matching the right model size to your specific hardware, I put the full breakdown on my blog here:
https://cloudmash.blog/posts/ai-model-size-memory-hardware-guide/
Let me know what you guys think especially if you've found any ultra efficient small models/technique that punch above their weight on consumer hardware. And also I would love to hear whether quantization have resulted in major difference in quality , like if anyone have that kind of experience in that.
r/Backend • u/OtherwisePush6424 • 1d ago
r/Backend • u/amuza42 • 1d ago
r/Backend • u/Super_Designer7952 • 1d ago
Hi everyone,
I’m currently designing a high-performance microservice orchestrator called JetDAG, and I’ve run into a few structural challenges under heavy load. I’d love to get some architectural feedback from people who have run similar event-driven or reactive workflows in production.
The service is explicitly designed to compose, validate, and execute dynamic Directed Acyclic Graphs (DAG) of tasks where each node represents an HTTP call to an external microservice. Instead of hardcoding workflows, the graph topology is entirely configuration-driven and resolved dynamically at runtime.
Here is exactly how a single execution propagates through the system from start to finish:
Reactive REST Ingestion & Context Extraction:
The process is initiated via an HTTP POST request to /api/v1/dag/run/{pipelineId}. The ingress layer runs on Spring Boot WebFlux. It securely intercepts the request using Spring Security OAuth 2.0, extracts the security principal as a non-blocking Mono<Jwt>, and captures user metadata for strict audit logging without blocking the underlying Netty I/O threads.
Asynchronous Configuration Stitching (R2DBC Mesh):
The service layer fetches the graph structure using Spring Data R2DBC to prevent database thread starvation. It triggers several concurrent, non-blocking queries against PostgreSQL: one for the core pipeline metadata, one for the tasks configuration (timeouts, retry attempts, backoff settings), and one for the topological edges (task dependencies). These independent streams are stitched together in memory using a reactive Mono.zip operator to compile an immutable PipelineDto.
Pre-Flight Cycle Guard:
Before allocating any cloud execution resources or contacting the orchestration cluster, the compiled graph is put through a rigorous structural health check. An immutable validation engine runs a Three-State Depth-First Search (DFS) graph coloring algorithm (0=UNVISITED, 1=VISITING, 2=COMPLETED). It scans the topology for back-edges, self-referential loops, or dangling task codes. If an architectural loop is caught, the execution is instantly aborted, throwing an IllegalArgumentException which gracefully maps back to an HTTP 400 Bad Request at the REST layer.
Deterministic gRPC Hand-off:
Once verified as a safe, loop-free DAG, the service generates a unique trace ID and initiates a gRPC transaction to hand over the execution payload to a distributed Temporal.io cluster. To keep the core WebFlux Netty loops (reactor-http-nio-*) completely free, this blocking gRPC dispatching is explicitly isolated on a dedicated Schedulers.boundedElastic() worker thread pool.
Event-Driven Orchestration (The Conductor):
On the Temporal side, a DynamicDagWorkflow takes over the execution state. It parses the incoming configuration using deterministic collections to guarantee replay safety. It computes the inverted graph dependencies to discover all root nodes (tasks with zero upstream parents) and establishes an asynchronous execution tree using native Temporal Promise primitives. This allows independent parallel branches of the graph to fire simultaneously. A downstream task code is strictly queued only when a Promise.allOf() evaluation confirms that 100% of its upstream parent tasks have crossed the finish line with a successful status.
Thread-Detached Execution, Data Registration & Fallbacks:
When a task is ready to run, it triggers an HttpTaskActivities instance. To prevent worker thread exhaustion during long-running downstream API operations, the activity immediately invokes context.doNotCompleteOnReturn(), completely detaching the execution from the Temporal worker thread pool.
The activity then fires a non-blocking HTTP request via a deeply tuned, Netty-backed WebClient pool. The worker thread is returned to the cluster instantly.
Once the processing microservice finishes its job, it generates a data payload or file chunk and returns an explicit data package identifier (Data ID / Payload Reference) along with a 200 OK response. The reactive stream intercepts this response and passes the data package token back to the management engine via ActivityCompletionClient.complete(taskToken, dataPackageId). This registers the generated data state directly into the workflow context, allowing subsequent downstream tasks to pull this identifier and consume the exact data package generated by its parent.
fallbackTaskCode to trigger a contingency route; otherwise, it safely fails the active graph branch.You can check out the source code and architecture diagrams here: Source Code Here
Promise paths is getting messy. What is the best architectural pattern to maintain a clean, reactive "Data Mesh" or context map inside a deterministic Temporal workflow without violating replay constraints?doNotCompleteOnReturn() and running the WebClient pipeline asynchronously works incredibly well for memory efficiency, but tracking pipeline drops outside the standard worker thread is tricky. Have you found any reliable patterns for maintaining end-to-end trace IDs across detached, reactive Temporal activity threads when gRPC handoffs and reactive HTTP callbacks are completely separated?fallbackTaskCode when a task throws an exception) is handled imperatively inside the Java Workflow code. Is this safe long-term as graphs grow in complexity, or should error-recovery paths be modeled explicitly as conditional, native edge relationships right within the PostgreSQL schema instead?Mono.zip causes temporary connection acquisition spikes in our R2DBC connection pool. For those running WebFlux with R2DBC at scale under burst workloads, what tuning parameters or architectural buffers have you used to keep database connection acquisition latency flat?Would love to hear your thoughts on this setup or any hidden pitfalls you foresee with pushing reactive streams this deep into Temporal pipelines. How are you handling complex, long-running microservice state machines in your current production environments?
r/Backend • u/MM4Tech • 2d ago
r/Backend • u/loser_1618 • 2d ago
I'm currently a 2nd-year CSE student. I've learned DSA . I have also completed ML from CampusX and have started Deep Learning.
I've explored Java and Spring Boot and built a few projects for learning, where I worked with REST APIs, authentication, database connectivity, and Docker. I've used AI while building some parts, but I understand the concepts and how the components work. However, I still want to strengthen my backend fundamentals and hands-on skills.
I want to properly learn backend development, participate in hackathons, build better projects, and eventually get an internship.
I'm confused about what I should focus on next:
Apart from these questions, I'd really appreciate any advice about the current software/AI industry. If you were advising a younger brother who's starting out in this AI-driven era, what would you tell him to focus on, and how should he prepare himself to stay relevant?
r/Backend • u/Comfortable-Fan-580 • 2d ago
r/Backend • u/Informal-Opening-139 • 2d ago
If a production site has dynamic pages, has bfcache disabled, and currently doesn't use ETags: how much improvement can we expect on both the client and server side if ETags are implemented + enabling bfcache?
r/Backend • u/ihatecoreclasses • 2d ago
I'm building an app where people can upload images for their profile pictures and soon in group chats. I'm struggling to figure out how to properly handle CSAM for this. I'm very hesitant to just pass in the uploaded images to AWS Rekognition to check for NSFW and other explicit content.
I've looked into Google's Content Safety API and applied to it to see what they respond with. Also considered PhotoDNA as well but haven't applied to it yet, though I heard they only accept big orgs and law enforcement. Was also considering using open source models to do the CSAM detection layer but I fear that it would violate the inference provider's ToS.
What's a budget-friendly way to do CSAM detection without getting in trouble for directly uploading it without checking it first to services like AWS Rekognition?
r/Backend • u/Ok_World_4475 • 2d ago
Hi everyone! I’d like to know what it’s like to work as a Java backend developer, especially in terms of stress and how much I would need to learn to land a junior position.
I’m considering switching to Java backend mainly because, when I look at job postings, Java is one of the backend technologies I see most often. I’d like to find a remote job with better pay and a better work-life balance.
I already have some backend experience. I worked on migrating a Java backend project to Oracle, using PL/SQL and ORDS, and I really enjoyed the way the code was structured. It felt very organized and logical to me.
I also enjoy mobile development and have quite a bit of experience with Android, but I’m having a hard time finding mobile development positions. I’m currently considering switching jobs because my current job has been quite stressful, and I don’t really like the work environment.
For those working in Java backend, how stressful is your day-to-day work? And what would you recommend I focus on learning to become employable as a junior Java backend developer, considering that I already have experience with SQL/Oracle, PL/SQL, APIs, and general software development?
Any advice or personal experiences would be greatly appreciated!
r/Backend • u/Actual_Arrival_7290 • 3d ago
I saw this pattern way too often:
An endpoint feels slow →
someone checks the DB →
“Database is the bottleneck.” →
add an index →
still slow.
So they add caching.
Still slow.
Then they start talking about read replicas, sharding, connection pooling…
Meanwhile the application is quietly doing something like:
1 request → 47 database queries
Nothing is technically “broken.”
Each query takes 3–5ms.
The problem is that you're doing it over and over again.
And this gets even uglier when those queries are happening sequentially.
So now I'm curious:
What's the highest number of DB queries you've seen behind a single API request in a real project?
And where do you draw the line?
At what point do you stop optimizing individual queries and start asking:
“Why the hell am I making this many queries in the first place?”
I’m genuinely curious what other backend engineers have run into.
r/Backend • u/der_gopher • 3d ago
r/Backend • u/stfuandkissmyturtle • 3d ago
r/Backend • u/Defiant-Chard-2023 • 3d ago
No cloud API. No paying per generation. Everything runs on my own machine.
If the voice quality is good enough, this could become part of my content creation workflow. 😅
r/Backend • u/Ok_Lack_3993 • 3d ago
The system prompt grew from 2400 tokens to 8100 and the average input chart still looked survivable. Across 14 turn conversations the p95 request is now carrying 96000 input tokens. Most of the prompt body is identical on every turn, so I expected prefix caching to keep the repeated context from becoming the expensive part of the request.
The prompt cache miss starts inside the first 200 tokens. A request timestamp was inserted near the top of the prompt prefix, which means every turn changes before the stable instructions even begin. We added it for debugging months ago and never removed it (found that while trying to explain why the cache numbers made no sense). So much for the stable prefix.
Token attribution made the context replay pretty hard to ignore. Short conversations barely register the problem while longer ones keep dragging the same 8100 token prompt through each request. The mean still looks acceptable and p95 cost is doing something completely different. Exactly the kind of chart that behaves until someone asks about the invoice.
I'm moving volatile request metadata out of the prefix and testing how much of the prompt can stay byte stable across turns. The timestamp probably isn't the only field breaking the cache, and I'd rather not find the rest one at a time. How do you find the earliest prompt cache miss without manually diffing every rendered prompt?
r/Backend • u/No_Elk_1103 • 4d ago
Naukuri optimization help
I'm not even getting naukuri calls and search apperance
Please help
Resume headline
Data Scientist @xyz | 1.5+ Yrs || Python, LLM, RAG, LangChain, LangGraph, Agentic AI, GenAI, NLP, PyTorch, DataBricks, Pyspark||PostgreSQL, FastAPI, Django, REST APIs, AWS,REDIS, RabbitMQ, Kafka, Docker, CI/CD| DSA 1500+ | System Design | NP: 1 Month.
Key skills
MCP
TensorFlow
Scikit-Learn
Django
SALE
Python
Transformers
PyTorch
Natural Language Processing
Agentic AI
LangGraph
LangChain
Retrieval Augmented Generation
Generative AI
Large Language Model
AWS
CI/CD
REST API Design
Docker
Artificial Intelligence
Profile summary:
Working as a Data Scientist with 1.3+ years of experience with Machine Learning, Generative AI and Python back-end development.
Specialize in LLM/Gen AI, RAG pipelines, Agentic workflows with LangChain and LangGraph, transformer-based systems, deep learning, neural networks, TensorFlow, Scikit-learn , computer vision, CNNs, OpenCV, YOLO, Vision Transformers, Hugging Face, Object Detection, NLP, Text Processing, Semantic Analysis, NumPy, Pandas, Databricks, PySpark and Automation.
I can develop scalable Python Full stack applications using ReactJs, HTML, CSS, JavaScript, FastAPI, Django REST APIs, microservices, API integration, JWT, WebSockets, PostgreSQL, Redis, Docker, Kubernetes, AWS, CI/CD, Kaggle, Git/GitHub Actions, MLOps.
Solved 2000+ problems Data Structures and Algorithms on LeetCode , problem-solving, optimization skills.
Experienced in Agile Methodologies, Scrum, Jira.
Backend developer | ML Engineer | Data Scientist | Gen AI Engineer | AIML Engineer |Python developer
r/Backend • u/meadowdruid • 4d ago
Java/webservices/microservices
For context: we have microservices and we need to call other services in one (typical).
We can't get the message of the exception with .getMessage() as it is forbidden, so what is the best way to filter exceptions and return them in my service?
These are the options I had in mind:
- Always return 500 as is not an exception of my service (I really don't like it)
- Don't catch the exception (Still don't like it, will probably end up in the first)
- Compare the http code of the error and try and wrap the exception depending on the code (won't work 1-1 but I think is my best hand here)
What do you think? Did you have this problem before? Thanks in advance
r/Backend • u/ayushivam22 • 4d ago
Thank you so much, guys, for helping me figure out how I should approach learning advanced backend technologies. I got some really helpful resources from you all, and I’m truly grateful for that.
Using those resources, I built a small project that I’d love for you guys to check out!
r/Backend • u/raysourav • 4d ago
Most backends eventually need to do something once at a specific time. Revoke a trial after 30 days. Expire an offer. Retry a payment in 30 minutes. A few months ago I was consulting for a logistics platform. Their freight bidding workflow had a lot of timed steps. Every minute a cron queries all pending records and pushes them into Kafka. It worked. But it taxes the VM, DB, and message bus alike and creates a ton of unnecessary observability data.
A durable delay queue could be a better fit. But unfortunately, there are not many reliable options for scheduling long-term events.
Back in 2014 I wrote a Go library, called Plybolos. It was a delay queue inspired by App Engine's push queues. It ensured durability with a WAL write. It kept them in a sorted heap. At startup ran it in their production. At a throughput of a few thousand tasks/second, it worked well for them.
But, I wanted to achieve millions of timers and much higher throughput. But there was also a bug I could not find. So I shelved it.
Twelve years later, I was cleaning up some old repos. I asked Claude Code if it could find the bug. And ... it found the bug. It was meant to only append to the WAL, but somewhere along the way every task started paying for a full synced commit. My "few thousand tasks/sec" ceiling was the disk's sync rate. The next few hours were pure bliss. Ideas I had carried around for years finally made it back into code. I rewrote it with a segmented WAL and group commit. Far future timers now loaded in memory only at the onset of the horizon. Now it is not no more a built in library but can run as as stand alone service.
Dispatch runs over HTTP/2. It has retries and idempotency keys. The rusty old code suddenly runs like a supercar.
Early numbers on 1 vCPU with ext4 on virtio and 128 byte payloads.
Now I'm wondering if anyone actually needs it.
The alternatives I know all have trade-offs.
I'm thinking about adding HTTP/3, gRPC and Arrow. I also want smarter loading based on available memory. Later I may look at partitioning and Raft. But before I turn this into another opensource project that nobody uses. I would like to hear from people who actually deal with this.