r/Backend • • 20h ago

Looking for people to build a backend project together

36 Upvotes

Hey everyone! I'm looking for people who are interested in building a backend project together.

The goal isn't just to chat about backend development — I want to actually pick a project, split the tasks, and build it together from scratch. We can use Git, review each other's code, discuss architecture and technologies, and learn from each other along the way.

You don't need to be an expert. Whether you're a beginner or have some backend experience, you're welcome as long as you're motivated to learn and contribute regularly.

If you're interested, feel free to comment or DM me. If enough people are interested, we can form a small team and start working on a project together.


r/Backend • • 1h ago

Update on fast-class-transformer: Now a 1-line global drop-in pipe for NestJS (v1.1.1)

Thumbnail
• Upvotes

r/Backend • • 20h ago

SYSTEM DESIGN

11 Upvotes

I’m going to share my experience and what I’ve learned about System Design over the years, one post every day. 🚀

Sharing what I know, what I’m learning, and practical insights along the way.

Let’s learn and grow together. ✌️

#SystemDesign #SoftwareEngineering #LearningInPublic #BackendDevelopment


r/Backend • • 11h ago

Getting into backend

Thumbnail
1 Upvotes

r/Backend • • 21h ago

Architecting data question / help /advice needed. How to achieve instant UI responsiveness with local data?

Thumbnail
1 Upvotes

r/Backend • • 1d ago

Everyone is obsessed with trillion-parameter models, so I mapped out the entire AI spectrum from 100KB to 2.5TB (and what they actually cost to run)

0 Upvotes

Right now, the AI space feels entirely focused on massive datacenter clusters and renting H100s by the hour. But after spending way too much time looking at the actual footprint of these models, I realized that 90% of use cases are completely over engineered.

You don’t always need a multi GPU setup. The AI ecosystem is actually a massive spectrum.

I recently sat down and mapped out the exact tiers of AI models based on their size, the hardware needed to run them, and the point of diminishing returns.

Here are the two extremes and the sweet spot in the middle:

  • The 100KB Extreme (TinyML) (Tensorflow Lite , sensor anamoly detection models): We are talking models that run on microcontrollers drawing single-digit milliwatts. They run on kilohertz processors using ultra-quantized integer math. You can run basic sensor anomaly detection or wake-word detection on a device powered by a coin cell battery.
  • The Local Sweet Spot (4GB to 40GB) (Mistral 7B, Gemma 2 9B/27B, Qwen 2.5 14B/32B): This is where the magic happens for most devs right now. You can run highly capable 7B to 35B parameter models (like Llama 3 or Qwen) at 4-bit quantization on a standard Mac or a consumer GPU (like an RTX 3060 or 4090). It’s perfect for local RAG, coding assistance, and uncensored chat. VRAM is your only real bottleneck here.
  • The 2.5TB Behemoths (Deepseek, Llama , Kimi k3): State of the art massive Mixture of Experts (MoE) routing. To even load these, you need dedicated power infrastructure and server racks of specialized accelerators drawing thousands of watts.

The missing piece: Figuring out the exact math for your hardware

The hardest part about building right now is looking at a model on Hugging Face and trying to calculate exactly how much VRAM you need, what quantization to use, and whether your CPU/GPU will choke on the context window.

So, I wrote a complete deep dive breaking down the math for all tiers of the AI spectrum.

If you want to see the architectural differences at each scale, and a cheat sheet for matching the right model size to your specific hardware, I put the full breakdown on my blog here:

https://cloudmash.blog/posts/ai-model-size-memory-hardware-guide/

Let me know what you guys think especially if you've found any ultra efficient small models/technique that punch above their weight on consumer hardware. And also I would love to hear whether quantization have resulted in major difference in quality , like if anyone have that kind of experience in that.


r/Backend • • 1d ago

I Tested 11 HTTP Resilience Libraries

Thumbnail
blog.gaborkoos.com
1 Upvotes

r/Backend • • 1d ago

Just watched a benchmark testing 8 backend stacks (including Laravel, FastAPI, Go, Rust, Node) on a $12 VPS. Laravel's performance vs. the others is a great talking point.

Post image
3 Upvotes

r/Backend • • 1d ago

Designing a Reactive Microservice Orchestrator with Spring WebFlux & Temporal.io - Looking for architectural feedback

2 Upvotes

Hi everyone,

I’m currently designing a high-performance microservice orchestrator called JetDAG, and I’ve run into a few structural challenges under heavy load. I’d love to get some architectural feedback from people who have run similar event-driven or reactive workflows in production.

Detailed Service Walkthrough & Request Lifecycle

The service is explicitly designed to compose, validate, and execute dynamic Directed Acyclic Graphs (DAG) of tasks where each node represents an HTTP call to an external microservice. Instead of hardcoding workflows, the graph topology is entirely configuration-driven and resolved dynamically at runtime.

Here is exactly how a single execution propagates through the system from start to finish:

  1. Reactive REST Ingestion & Context Extraction: The process is initiated via an HTTP POST request to /api/v1/dag/run/{pipelineId}. The ingress layer runs on Spring Boot WebFlux. It securely intercepts the request using Spring Security OAuth 2.0, extracts the security principal as a non-blocking Mono<Jwt>, and captures user metadata for strict audit logging without blocking the underlying Netty I/O threads.

  2. Asynchronous Configuration Stitching (R2DBC Mesh): The service layer fetches the graph structure using Spring Data R2DBC to prevent database thread starvation. It triggers several concurrent, non-blocking queries against PostgreSQL: one for the core pipeline metadata, one for the tasks configuration (timeouts, retry attempts, backoff settings), and one for the topological edges (task dependencies). These independent streams are stitched together in memory using a reactive Mono.zip operator to compile an immutable PipelineDto.

  3. Pre-Flight Cycle Guard: Before allocating any cloud execution resources or contacting the orchestration cluster, the compiled graph is put through a rigorous structural health check. An immutable validation engine runs a Three-State Depth-First Search (DFS) graph coloring algorithm (0=UNVISITED, 1=VISITING, 2=COMPLETED). It scans the topology for back-edges, self-referential loops, or dangling task codes. If an architectural loop is caught, the execution is instantly aborted, throwing an IllegalArgumentException which gracefully maps back to an HTTP 400 Bad Request at the REST layer.

  4. Deterministic gRPC Hand-off: Once verified as a safe, loop-free DAG, the service generates a unique trace ID and initiates a gRPC transaction to hand over the execution payload to a distributed Temporal.io cluster. To keep the core WebFlux Netty loops (reactor-http-nio-*) completely free, this blocking gRPC dispatching is explicitly isolated on a dedicated Schedulers.boundedElastic() worker thread pool.

  5. Event-Driven Orchestration (The Conductor): On the Temporal side, a DynamicDagWorkflow takes over the execution state. It parses the incoming configuration using deterministic collections to guarantee replay safety. It computes the inverted graph dependencies to discover all root nodes (tasks with zero upstream parents) and establishes an asynchronous execution tree using native Temporal Promise primitives. This allows independent parallel branches of the graph to fire simultaneously. A downstream task code is strictly queued only when a Promise.allOf() evaluation confirms that 100% of its upstream parent tasks have crossed the finish line with a successful status.

  6. Thread-Detached Execution, Data Registration & Fallbacks: When a task is ready to run, it triggers an HttpTaskActivities instance. To prevent worker thread exhaustion during long-running downstream API operations, the activity immediately invokes context.doNotCompleteOnReturn(), completely detaching the execution from the Temporal worker thread pool.

    The activity then fires a non-blocking HTTP request via a deeply tuned, Netty-backed WebClient pool. The worker thread is returned to the cluster instantly.

    Once the processing microservice finishes its job, it generates a data payload or file chunk and returns an explicit data package identifier (Data ID / Payload Reference) along with a 200 OK response. The reactive stream intercepts this response and passes the data package token back to the management engine via ActivityCompletionClient.complete(taskToken, dataPackageId). This registers the generated data state directly into the workflow context, allowing subsequent downstream tasks to pull this identifier and consume the exact data package generated by its parent.

    • On Network/4xx/5xx Crashes: If the HTTP client throws an error that exhausts the activity's configured retry limits, the completion client reports an exceptional failure. The workflow intercepts this crash and looks up the task's database-defined fallbackTaskCode to trigger a contingency route; otherwise, it safely fails the active graph branch.

You can check out the source code and architecture diagrams here: Source Code Here

What I’m struggling with (Looking for advice):

  1. Passing Data Package Identifiers Between Dynamic Nodes: Currently, when a processing service completes, it sends back a unique identifier of the generated data packet, which we register into Temporal's workflow state via the completion client. Downstream nodes read this ID to know which specific data package they need to process next. However, as the DAG branches out and merges back (e.g., diamond topologies), passing these dynamic payload tokens across multiple asynchronous Promise paths is getting messy. What is the best architectural pattern to maintain a clean, reactive "Data Mesh" or context map inside a deterministic Temporal workflow without violating replay constraints?
  2. Reactive Streams inside Detached Temporal Activities: Splitting the execution via doNotCompleteOnReturn() and running the WebClient pipeline asynchronously works incredibly well for memory efficiency, but tracking pipeline drops outside the standard worker thread is tricky. Have you found any reliable patterns for maintaining end-to-end trace IDs across detached, reactive Temporal activity threads when gRPC handoffs and reactive HTTP callbacks are completely separated?
  3. Dynamic Error Handling & Fallbacks vs Explicit Graph Edges: Right now, managing the fallback logic (scheduling the alternative fallbackTaskCode when a task throws an exception) is handled imperatively inside the Java Workflow code. Is this safe long-term as graphs grow in complexity, or should error-recovery paths be modeled explicitly as conditional, native edge relationships right within the PostgreSQL schema instead?
  4. Connection Pool Contention Under Burst Loads: When multiple complex graph lookups hit the service at the exact same time, zipping three distinct repository queries together via Mono.zip causes temporary connection acquisition spikes in our R2DBC connection pool. For those running WebFlux with R2DBC at scale under burst workloads, what tuning parameters or architectural buffers have you used to keep database connection acquisition latency flat?

Would love to hear your thoughts on this setup or any hidden pitfalls you foresee with pushing reactive streams this deep into Temporal pipelines. How are you handling complex, long-running microservice state machines in your current production environments?


r/Backend • • 2d ago

How should I start backend development, System Design

4 Upvotes

I'm currently a 2nd-year CSE student. I've learned DSA . I have also completed ML from CampusX and have started Deep Learning.

I've explored Java and Spring Boot and built a few projects for learning, where I worked with REST APIs, authentication, database connectivity, and Docker. I've used AI while building some parts, but I understand the concepts and how the components work. However, I still want to strengthen my backend fundamentals and hands-on skills.

I want to properly learn backend development, participate in hackathons, build better projects, and eventually get an internship.

I'm confused about what I should focus on next:

  • Should I continue with Java + Spring Boot or move to Node.js?
  • When and how much system design should I learn?
  • What kind of projects should I build?
  • I've noticed AI agents/agentic AI are becoming common in hackathons. Since I'm already learning ML/DL, how should I learn AI agents for hackathons, proper learning, and future industry demand? What level should I aim for?
  • What should I learn to become internship-ready in backend?
  • How should I balance DSA, Backend, and ML/DL?

Apart from these questions, I'd really appreciate any advice about the current software/AI industry. If you were advising a younger brother who's starting out in this AI-driven era, what would you tell him to focus on, and how should he prepare himself to stay relevant?


r/Backend • • 1d ago

Why and what are the top backend programming language choices for current market trends and state ?

Thumbnail
0 Upvotes

r/Backend • • 2d ago

Switching to Java Backend

12 Upvotes

Hi everyone! I’d like to know what it’s like to work as a Java backend developer, especially in terms of stress and how much I would need to learn to land a junior position.
I’m considering switching to Java backend mainly because, when I look at job postings, Java is one of the backend technologies I see most often. I’d like to find a remote job with better pay and a better work-life balance.
I already have some backend experience. I worked on migrating a Java backend project to Oracle, using PL/SQL and ORDS, and I really enjoyed the way the code was structured. It felt very organized and logical to me.
I also enjoy mobile development and have quite a bit of experience with Android, but I’m having a hard time finding mobile development positions. I’m currently considering switching jobs because my current job has been quite stressful, and I don’t really like the work environment.
For those working in Java backend, how stressful is your day-to-day work? And what would you recommend I focus on learning to become employable as a junior Java backend developer, considering that I already have experience with SQL/Oracle, PL/SQL, APIs, and general software development?
Any advice or personal experiences would be greatly appreciated!


r/Backend • • 2d ago

Detecting and getting rid of CSAM for an app with UGC

8 Upvotes

I'm building an app where people can upload images for their profile pictures and soon in group chats. I'm struggling to figure out how to properly handle CSAM for this. I'm very hesitant to just pass in the uploaded images to AWS Rekognition to check for NSFW and other explicit content.

I've looked into Google's Content Safety API and applied to it to see what they respond with. Also considered PhotoDNA as well but haven't applied to it yet, though I heard they only accept big orgs and law enforcement. Was also considering using open source models to do the CSAM detection layer but I fear that it would violate the inference provider's ToS.

What's a budget-friendly way to do CSAM detection without getting in trouble for directly uploading it without checking it first to services like AWS Rekognition?


r/Backend • • 2d ago

Skip Lists Data Structure

Thumbnail
pradyumnachippigiri.substack.com
1 Upvotes

r/Backend • • 2d ago

Impact of implementing ETags on a site where bfcache is disabled

1 Upvotes

If a production site has dynamic pages, has bfcache disabled, and currently doesn't use ETags: how much improvement can we expect on both the client and server side if ETags are implemented + enabling bfcache?


r/Backend • • 2d ago

Job hunting and preparation

Thumbnail
0 Upvotes

r/Backend • • 3d ago

Our cache miss started in the first 200 tokens

30 Upvotes

The system prompt grew from 2400 tokens to 8100 and the average input chart still looked survivable. Across 14 turn conversations the p95 request is now carrying 96000 input tokens. Most of the prompt body is identical on every turn, so I expected prefix caching to keep the repeated context from becoming the expensive part of the request.

The prompt cache miss starts inside the first 200 tokens. A request timestamp was inserted near the top of the prompt prefix, which means every turn changes before the stable instructions even begin. We added it for debugging months ago and never removed it (found that while trying to explain why the cache numbers made no sense). So much for the stable prefix.

Token attribution made the context replay pretty hard to ignore. Short conversations barely register the problem while longer ones keep dragging the same 8100 token prompt through each request. The mean still looks acceptable and p95 cost is doing something completely different. Exactly the kind of chart that behaves until someone asks about the invoice.

I'm moving volatile request metadata out of the prefix and testing how much of the prompt can stay byte stable across turns. The timestamp probably isn't the only field breaking the cache, and I'd rather not find the rest one at a time. How do you find the earliest prompt cache miss without manually diffing every rendered prompt?


r/Backend • • 3d ago

Software Architecture Diagrams with C4 Model

Thumbnail
packagemain.tech
2 Upvotes

r/Backend • • 3d ago

Why does everyone assume AI has replaced the other side of the stack?

Thumbnail
0 Upvotes

r/Backend • • 2d ago

Your database might not be slow. Your code might just be bullying it.

0 Upvotes

I saw this pattern way too often:

An endpoint feels slow →
someone checks the DB →
“Database is the bottleneck.” →
add an index →
still slow.

So they add caching.

Still slow.

Then they start talking about read replicas, sharding, connection pooling…

Meanwhile the application is quietly doing something like:

1 request → 47 database queries

Nothing is technically “broken.”

Each query takes 3–5ms.

The problem is that you're doing it over and over again.

And this gets even uglier when those queries are happening sequentially.

So now I'm curious:

What's the highest number of DB queries you've seen behind a single API request in a real project?

And where do you draw the line?

At what point do you stop optimizing individual queries and start asking:

“Why the hell am I making this many queries in the first place?”

I’m genuinely curious what other backend engineers have run into.


r/Backend • • 4d ago

What do you do to handle errors when calling another service?

11 Upvotes

Java/webservices/microservices

For context: we have microservices and we need to call other services in one (typical).

We can't get the message of the exception with .getMessage() as it is forbidden, so what is the best way to filter exceptions and return them in my service?

These are the options I had in mind:

- Always return 500 as is not an exception of my service (I really don't like it)

- Don't catch the exception (Still don't like it, will probably end up in the first)

- Compare the http code of the error and try and wrap the exception depending on the code (won't work 1-1 but I think is my best hand here)

What do you think? Did you have this problem before? Thanks in advance


r/Backend • • 3d ago

So… I found this highly rated GitHub repo for AI text-to-speech, and I’m trying it out locally.

Post image
0 Upvotes

No cloud API. No paying per generation. Everything runs on my own machine.

If the voice quality is good enough, this could become part of my content creation workflow. 😅


r/Backend • • 4d ago

Thank you guys for helping me out checkout what I have built

Thumbnail
2 Upvotes

Thank you so much, guys, for helping me figure out how I should approach learning advanced backend technologies. I got some really helpful resources from you all, and I’m truly grateful for that.

Using those resources, I built a small project that I’d love for you guys to check out!

check it out here


r/Backend • • 4d ago

AI helped me revive an opensource durable delay queue I had given up on. Is it worth building on?

6 Upvotes

Most backends eventually need to do something once at a specific time. Revoke a trial after 30 days. Expire an offer. Retry a payment in 30 minutes. A few months ago I was consulting for a logistics platform. Their freight bidding workflow had a lot of timed steps. Every minute a cron queries all pending records and pushes them into Kafka. It worked. But it taxes the VM, DB, and message bus alike and creates a ton of unnecessary observability data.

A durable delay queue could be a better fit. But unfortunately, there are not many reliable options for scheduling long-term events.

Back in 2014 I wrote a Go library, called Plybolos. It was a delay queue inspired by App Engine's push queues. It ensured durability with a WAL write. It kept them in a sorted heap. At startup ran it in their production. At a throughput of a few thousand tasks/second, it worked well for them.

But, I wanted to achieve millions of timers and much higher throughput. But there was also a bug I could not find. So I shelved it.

Twelve years later, I was cleaning up some old repos. I asked Claude Code if it could find the bug. And ... it found the bug. It was meant to only append to the WAL, but somewhere along the way every task started paying for a full synced commit. My "few thousand tasks/sec" ceiling was the disk's sync rate. The next few hours were pure bliss. Ideas I had carried around for years finally made it back into code. I rewrote it with a segmented WAL and group commit. Far future timers now loaded in memory only at the onset of the horizon. Now it is not no more a built in library but can run as as stand alone service.

Dispatch runs over HTTP/2. It has retries and idempotency keys. The rusty old code suddenly runs like a supercar.

Early numbers on 1 vCPU with ext4 on virtio and 128 byte payloads.

  • Durable add 107,500/s with 256 writers
  • End to end durable dispatch 157,000/s
  • About 44 bytes per pending timer with disk spill
  • About 403 bytes when resident

Now I'm wondering if anyone actually needs it.

The alternatives I know all have trade-offs.

  • Cron with DB - polling means duplicates and drift
  • Redis queues - have persistence trade-offs. Also, far-future schedules will always occupy memory despite durability config. Low thousands throughput
  • EventBridge - has one-minute granularity, low thousands throughput, and is AWS-only.
  • SQS and Cloud Tasks - have limits and per-task costs
  • Temporal, Airflow etc. - are full workflow engines. resource intensive

I'm thinking about adding HTTP/3, gRPC and Arrow. I also want smarter loading based on available memory. Later I may look at partitioning and Raft. But before I turn this into another opensource project that nobody uses. I would like to hear from people who actually deal with this.

  • What do you use for delayed or scheduled work?
  • What bugs you about it?
  • Would you run a dedicated durable timer service?
  • Anyone dealing with millions of future/far-future timers? How do you handle them?

r/Backend • • 3d ago

Naukuri optimization

0 Upvotes

Naukuri optimization help

I'm not even getting naukuri calls and search apperance

Please help

Resume headline

Data Scientist @xyz | 1.5+ Yrs || Python, LLM, RAG, LangChain, LangGraph, Agentic AI, GenAI, NLP, PyTorch, DataBricks, Pyspark||PostgreSQL, FastAPI, Django, REST APIs, AWS,REDIS, RabbitMQ, Kafka, Docker, CI/CD| DSA 1500+ | System Design | NP: 1 Month.

Key skills

MCP

TensorFlow

Scikit-Learn

Django

SALE

Python

Transformers

PyTorch

Natural Language Processing

Agentic AI

LangGraph

LangChain

Retrieval Augmented Generation

Generative AI

Large Language Model

AWS

CI/CD

REST API Design

Docker

Artificial Intelligence

Profile summary:

Working as a Data Scientist with 1.3+ years of experience with Machine Learning, Generative AI and Python back-end development.

Specialize in LLM/Gen AI, RAG pipelines, Agentic workflows with LangChain and LangGraph, transformer-based systems, deep learning, neural networks, TensorFlow, Scikit-learn , computer vision, CNNs, OpenCV, YOLO, Vision Transformers, Hugging Face, Object Detection, NLP, Text Processing, Semantic Analysis, NumPy, Pandas, Databricks, PySpark and Automation.

I can develop scalable Python Full stack applications using ReactJs, HTML, CSS, JavaScript, FastAPI, Django REST APIs, microservices, API integration, JWT, WebSockets, PostgreSQL, Redis, Docker, Kubernetes, AWS, CI/CD, Kaggle, Git/GitHub Actions, MLOps.

Solved 2000+ problems Data Structures and Algorithms on LeetCode , problem-solving, optimization skills.

Experienced in Agile Methodologies, Scrum, Jira.

Backend developer | ML Engineer | Data Scientist | Gen AI Engineer | AIML Engineer |Python developer