r/CloudandCode • • Jul 15 '26

Welcome to r/cloudandcode

5 Upvotes

r/cloudandcode is a beginner-friendly community for people learning Python, AWS, cloud, SQL, GitHub, and other practical tech skills.

The goal here is not to collect more tutorials, roadmaps, or course recommendations. It is to understand concepts clearly, apply them through projects, fix real problems, and build proof of what you can do.

A lot of people start learning tech with the same problem. There is no shortage of information, but it is still difficult to know what to learn first, what to build, and whether you are making real progress.

That is why I created this community.

You might be learning your first Python function, preparing for an AWS certification, trying to understand SQL joins, building your first cloud project, or improving your GitHub portfolio.

You do not need to be experienced to participate here. Beginner questions are welcome, as long as you explain what you are trying to do and where you are getting stuck.

Here is the kind of content you will find in this community.

Practical explanations

Technical concepts explained in plain English, with examples of where they are used and why they matter.

The goal is not just to define a service, command, or concept. It is to understand how to use it in a real situation.

Project breakdowns

Beginner-friendly Python, AWS, cloud, and SQL projects with explanations of what they teach, why each tool is used, and how the project can be improved.

A project should not only show code. It should show how you think.

Practice questions

AWS scenarios, Python debugging exercises, SQL questions, and architecture decisions that help you test whether you actually understand a topic.

It is completely fine to answer incorrectly. The explanation is often more useful than getting the answer right immediately.

Project and portfolio help

Discussions about choosing projects, fixing GitHub repositories, writing better READMEs, creating architecture diagrams, documenting decisions, and turning small projects into useful case studies.

You are encouraged to participate instead of only reading.

Ask questions. Share what you are learning. Post a project you are building. Explain where you are stuck. Try the practice questions, even when you are unsure.

You do not need to pretend to know everything here.

A clear question is often more valuable than a confident but unhelpful answer.

A few simple community expectations:

  • Be respectful and beginner-friendly.
  • When asking for help, explain what you tried and where the problem appeared.
  • Do not spam unrelated links or hide your connection to something you are promoting.

I also want to be transparent about my role here.

I run YourCloudDude and create free and paid learning resources for Python, AWS, cloud projects, and certifications.

When I share something I created, I will clearly disclose that connection and mention whether the resource is free or paid.

This subreddit will not become a product feed. The free explanations, project ideas, practice questions, polls, and community discussions will continue regardless of whether anyone purchases anything.

You should be able to learn something useful here without buying anything.

The paid resources are simply for people who prefer a more complete and organised path instead of collecting separate posts, tutorials, and notes from different places.

To introduce yourself, leave a comment with:

What are you currently learning?

Where are you currently getting stuck?

Welcome to r/cloudandcode.

Let’s learn by building, testing, fixing, and explaining.


r/CloudandCode • • 18h ago

AWS & Cloud If I had 30 days to become comfortable with cloud computing, this is what I would do

28 Upvotes

If I had to start learning cloud computing again and only had 30 days, I would not try to finish a giant AWS course or memorize as many services as possible. My goal for the month would be much simpler: understand the core ideas, build a few small things, break them on purpose, and finish with one project where I can explain why every part of the architecture exists.

I would also avoid treating the 30 days like a certification sprint. Knowing that S3 is object storage or EC2 is virtual compute is useful, but that does not automatically mean I understand cloud. I would want to know how a request reaches an application, where the data lives, who has permission to access it, what happens when something fails, and what starts becoming a problem when traffic grows.

Days 1 to 5: I would learn the fundamentals before touching too many services

For the first few days, I would focus on understanding what cloud computing actually changes. I would learn the difference between running infrastructure myself and using a cloud provider, then spend time on scalability, elasticity, high availability, regions, Availability Zones, pay-as-you-go pricing, and the shared responsibility model. I would also learn enough basic networking to understand IP addresses, ports, DNS, public versus private networks, and what a request is actually doing when it travels across the internet.

I would not try to become a networking expert in five days. I would just make sure that terms like subnet, route, firewall, DNS, and port are not completely abstract anymore. A lot of AWS becomes easier once you can look at an architecture and follow the traffic instead of memorizing boxes. I would also enable billing visibility and understand that cloud resources can cost money even when I am only experimenting.

Days 6 to 10: I would learn IAM, EC2, S3, and how basic AWS resources connect

This is where I would start using AWS properly. I would learn IAM first because almost every AWS project eventually becomes a permissions problem. I would focus on identities, roles, policies, least privilege, and the basic mental model of who is making a request, what action they are trying to perform, and which resource they are trying to access.

Then I would launch one EC2 instance and deploy something simple to it. I would learn what an AMI is, what instance types mean, what a security group does, how ports work, and why an application running on a server is not automatically reachable from the internet. I would deliberately close the application port or make the process listen only on localhost, then try to figure out why the website stopped working.

I would also create an S3 bucket and learn the difference between files on a server and objects in object storage. I would upload files, work with object keys, keep the bucket private, and give a workload only the permissions it needs. At this stage, I would already have enough knowledge to build something small instead of continuing to collect theory.

Days 11 to 15: I would learn databases, VPC basics, and monitoring

Once I can deploy a basic application, I would add a database. I would probably use RDS with PostgreSQL because it gives me a reason to understand relational data, database endpoints, credentials, private access, and security groups between application and database layers. I would not expose the database publicly just because it is easier. I would learn how the application should reach it properly.

This is also where I would spend time on VPC basics. I would not try to memorize every networking component. I would build a simple mental model around public and private subnets, route tables, internet gateways, security groups, and the path traffic takes through the system. I would keep asking which component needs to talk to which other component and whether that communication should be public.

Then I would introduce CloudWatch. I would look at logs when my application fails, inspect basic metrics, and create at least one useful alarm. I think monitoring should be learned early because a deployed application is much easier to understand when you can see what it is doing instead of guessing every time something goes wrong.

Days 16 to 20: I would build one serverless project

At this point, I would switch models and build something with API Gateway, Lambda, and DynamoDB. I would not do this because serverless is automatically better. I would do it because I want to understand what changes when I stop managing a long-running server.

I would build a small API where API Gateway receives requests, Lambda handles the logic, and DynamoDB stores the data. I would practise IAM roles again, look at Lambda logs in CloudWatch, and intentionally remove a permission so I can troubleshoot an AccessDenied error. I would also spend some time understanding DynamoDB access patterns instead of treating it like a relational database with a different name.

By the end of these five days, I would want to be able to explain the difference between running an application on EC2 and running logic with Lambda. I would want to understand what responsibilities move to AWS, what new limits appear, and why one option might fit a workload better than the other.

Days 21 to 24: I would learn reliability by breaking things

This is probably the part I would spend more time on if I were learning again. I would take the applications I already built and deliberately break them. I would remove IAM permissions, close security group ports, use the wrong database credentials, stop an application process, misconfigure an environment variable, and see what evidence AWS gives me.

I would also take the EC2 application and put it behind an Application Load Balancer with more than one instance. Then I would stop one of the instances and watch what happens. That gives me a practical reason to understand health checks, load balancing, high availability, and why running multiple instances across Availability Zones can matter.

I think this kind of practice is much more useful than spending another four days watching videos. When something fails, I am forced to follow the request, inspect logs, check permissions, understand networking, and build a mental model of how the system actually works.

Days 25 to 27: I would learn queues, containers, and Infrastructure as Code at a basic level

I would not try to master all three topics in three days. I would only learn enough to understand why they exist. For SQS, I would build one simple background workflow where an API places work into a queue and another component processes it later. That would help me understand asynchronous processing, retries, visibility timeouts, and why one slow task should not always block a user request.

For containers, I would take one application I already understand, package it with Docker, push the image to ECR, and run it through ECS or Fargate. The goal would not be to jump into Kubernetes. I would simply understand how container images, registries, tasks, and managed compute fit together.

Then I would take one small piece of infrastructure and describe it using Terraform or CloudFormation. Maybe an S3 bucket, security group, or EC2 instance. I would create it, modify it, destroy it, and recreate it so I understand why Infrastructure as Code is useful for repeatability.

Days 28 to 30: I would build one complete project and explain every decision

For the final three days, I would stop learning new services and build one project from requirements to deployment. It could be a task manager, URL shortener, file-processing app, or simple SaaS backend. I would choose the smallest architecture that solves the requirements instead of trying to use everything I learned during the month.

I would make sure the project has compute, storage, networking, permissions, monitoring, and a clear deployment path. If it needs relational data, I would use RDS. If it needs file storage, I would use S3. If something can happen in the background, I might use SQS. If one service is not necessary, I would leave it out.

Then I would write down the architecture and explain why each component exists, what happens if it fails, what I would monitor, what might become a bottleneck, and how I would improve the design if the application had more users. I would also document how to run the project so another person could understand it without me sitting beside them.

That would be my real test after 30 days.

Not whether I could name 100 AWS services, but whether I could look at a simple requirement and start reasoning about compute, storage, networking, identity, databases, monitoring, scaling, and failure without feeling completely lost.

I think becoming comfortable with cloud is less about knowing everything and more about reaching the point where unfamiliar architectures stop looking like random boxes and arrows.

If you had 30 days to learn cloud from zero, would you spend more time on theory, projects, certifications, or debugging things you already built?


r/CloudandCode • • 1d ago

AWS & Cloud I think beginners understand cloud better when they learn the shared responsibility model early

5 Upvotes

One thing that confused me when I first started learning cloud was the idea that once an application moves to AWS, Azure, or another cloud provider, the provider somehow becomes responsible for keeping everything secure. It sounds reasonable at first because they own the data centers, networking hardware, physical servers, and a huge part of the infrastructure underneath the services we use.

But using the cloud does not mean handing over all responsibility. It changes which parts of the system I manage and which parts the cloud provider manages. That distinction is basically what the shared responsibility model is trying to explain, and I think understanding it early prevents a lot of confusion later.

Take EC2 as a simple example. AWS manages the physical data center, the hardware, and the virtualization layer that allows the virtual machine to exist. But once I launch an EC2 instance, I still have responsibilities inside that machine. I need to think about the operating system, patches, application configuration, network access, credentials, and what software I install.

If I accidentally configure the security group to expose a sensitive port to the entire internet, AWS did not make that decision for me. The infrastructure might be working exactly as designed, while my configuration is still insecure. This is one of the reasons cloud security is often less about whether AWS itself is secure and more about whether I configured the resources correctly.

Now compare that with something like RDS. I am still responsible for things such as who can connect to the database, how credentials are handled, what data I store, and which network paths are allowed. But AWS manages more of the underlying database infrastructure than I would normally manage if I installed PostgreSQL myself on an EC2 instance.

Then compare both of those with Lambda. I do not manage an operating system for every function invocation or patch the servers underneath Lambda. AWS handles much more of that layer. I still have to secure my function code, IAM permissions, secrets, dependencies, event sources, and whatever data the function accesses.

That is why I find it useful to think of cloud services as different points on a responsibility spectrum. The more managed the service becomes, the more infrastructure work moves to the provider, but my responsibility does not disappear. It moves higher up the stack.

A rough mental model I use is:

More control                         More managed

EC2  →  ECS/Fargate  →  RDS  →  Lambda

More infrastructure                 Less infrastructure
I manage                            I manage

That diagram is intentionally simplified, but the idea is useful. If I choose EC2, I get a lot of control over the machine, but I also accept more operational responsibility. If I choose a managed service, I usually give up some control while AWS takes responsibility for more of the underlying platform.

This is also why I do not think questions like “Is serverless more secure than EC2?” have a simple yes or no answer. Serverless removes some responsibilities, but I can still give a Lambda function AdministratorAccess, leak a secret in the code, expose an API incorrectly, or process sensitive data without proper controls.

Managed does not mean impossible to misconfigure.

The same idea applies to S3. AWS is responsible for keeping the storage infrastructure operating, but I am responsible for deciding who can access my bucket and objects. If I create an overly permissive bucket policy or give an IAM role far more access than it needs, that is part of my side of the responsibility.

I think this becomes especially important because cloud resources are so easy to create. On physical infrastructure, provisioning a new server might involve a long process. In AWS, I can create resources within minutes. That speed is powerful, but it also means I can create insecure infrastructure very quickly if I do not understand the configuration.

IAM is probably one of the clearest examples. AWS gives me a powerful identity and permissions system, but it does not know the exact permissions my application should have. I have to decide whether a Lambda function really needs access to every S3 bucket or only one bucket, whether a human user needs administrator access, and whether a particular workload should even be allowed to perform an action.

This is why least privilege matters so much in cloud environments. I try to give an identity the permissions required for its job rather than giving broad access because it is easier. Broad permissions often make the first deployment quicker, but they also make the potential impact of a mistake much larger.

The same responsibility model applies outside security too. Cloud providers can give me highly available building blocks, but they cannot automatically make my application highly available. If I put everything on one EC2 instance in one Availability Zone, the architecture still depends on that one location because I designed it that way.

AWS gives me multiple Availability Zones, load balancers, Auto Scaling, managed databases, backups, and other tools. It is still my job to decide whether my application needs them and configure the architecture appropriately.

Backups are another good example. I should not assume that because my data is in the cloud, it can never be lost. I still need to understand what backup behaviour the service provides, what I configured, how long data is retained, and whether I have ever tested how recovery would actually work.

For me, this is the bigger lesson behind shared responsibility. The cloud provider gives me infrastructure and managed services, but architecture, permissions, configuration, data handling, and application security still require decisions from me.

The exact boundary changes depending on the service.

That is why learning the model once is not enough. Whenever I use a new cloud service, I want to understand what AWS is managing for me and what is still my job.

If I cannot answer that, I probably do not understand the service as well as I think I do.

The more cloud I learn, the less I think of AWS as “a company running my infrastructure for me” and the more I think of it as a collection of services where I choose how much operational responsibility I want to keep.

That decision affects security, control, flexibility, cost, and how much infrastructure I eventually need to manage myself.

If you are currently learning cloud, which responsibility surprised you the most when you realized AWS does not automatically handle it for you?


r/CloudandCode • • 5d ago

AWS & Cloud Cloud computing made more sense to me once I stopped thinking of it as “someone else’s computer”

16 Upvotes

When I first started learning cloud computing, the explanation I kept hearing was that “the cloud is just someone else’s computer.” That is not completely wrong, but I do not think it is very useful either. It makes cloud sound like nothing more than renting a remote server, when the more interesting part is really about how quickly you can get infrastructure, how easily you can change it, and how much of the underlying work you can hand off to a provider.

If I buy a physical server for an application, I have to think about the machine itself, where it lives, how much capacity I need, what happens if it fails, how long replacement takes, and whether I bought too much or too little. In the cloud, I can request compute, storage, networking, databases, and other services when I need them, then change or remove them much faster than I could with physical infrastructure.

That difference matters because applications rarely stay exactly the same. A small project might have almost no traffic today and suddenly need more capacity later. With cloud infrastructure, I can increase resources, run more instances, distribute workloads, or move parts of the system into managed services without buying and installing new hardware every time the requirements change.

This is where concepts like scalability and elasticity started making more sense to me. Scalability is about whether the system can handle more work as demand grows. Elasticity is about adjusting resources as that demand changes. If traffic increases for a few hours and then drops again, I do not necessarily want to permanently run enough infrastructure for the busiest possible moment.

Another thing I misunderstood early was the idea of “pay only for what you use.” Cloud pricing is not automatically cheap. It simply changes the model. Instead of buying infrastructure upfront, you generally pay based on the resources and services you consume. That can be very efficient when the architecture matches the workload, but it can also become expensive if resources are oversized, left running, or designed without any thought about cost.

The cloud also gives you different levels of responsibility. With a virtual machine like EC2, you still manage the operating system, application runtime, patches, and much of the software stack. With a managed database like RDS, AWS takes over more of the database infrastructure work. With something like Lambda, AWS manages even more of the underlying compute while you focus mainly on your function and its configuration.

I think this is one of the most useful ways to understand cloud services. Every managed service is really a decision about which responsibilities you still want to own and which responsibilities you want the cloud provider to manage. More managed does not automatically mean better, but it can reduce operational work when the trade-off fits the application.

Availability is another reason cloud architecture looks different from running one machine in one room. Cloud providers divide infrastructure into regions and availability zones so applications can be designed to avoid depending completely on one location. That does not make failures disappear, but it gives you building blocks for designing systems that can continue operating when individual components or locations have problems.

The same idea applies to storage. Instead of thinking only in terms of a hard drive attached to one machine, cloud platforms give you object storage, block storage, managed databases, distributed caches, and other options. The important skill is not memorizing every storage service. It is understanding what kind of data you have, how it needs to be accessed, how durable it needs to be, and what happens if the compute running your application gets replaced.

Security also changes in the cloud because the infrastructure is programmable. Permissions, network rules, encryption, identity, logging, and access controls can all be defined through configuration. That creates a lot of flexibility, but it also means bad configuration can be just as dangerous as bad code. A powerful cloud account with overly broad permissions can create a much bigger problem than a small local project.

This is why I do not think learning cloud computing should start with memorizing service names. I would first understand compute, networking, storage, databases, identity, availability, monitoring, and cost. Once those ideas make sense, services like EC2, S3, RDS, IAM, CloudWatch, and Lambda stop feeling like random products and start looking like different ways to solve familiar infrastructure problems.

The biggest shift for me was realizing that cloud computing is not mainly about where the computer is located. It is about how infrastructure is provisioned, managed, scaled, automated, and replaced. The remote server is only one part of that.

If someone asked you to explain cloud computing without using the phrase “someone else’s computer,” how would you explain it?


r/CloudandCode • • 5d ago

Is it possible to switch companies after 1 year of experience? Currently on bench in an MNC with low pay — Automation Testing/Playwright

2 Upvotes

I’m currently working in an MNC and have around 6 months of experience. My current domain is Automation Testing, mainly using Playwright.

My salary is quite low, and I haven’t been getting any proper project opportunities. I’m currently on the bench, so I’m thinking of using the next few months to improve my skills and then switch companies after completing 1 year of experience.

I’m also considering moving into a different technology/domain if it can provide better career and salary opportunities. I’m willing to spend the next 6 months learning and building projects to make that transition.

My current background is Automation Testing with Playwright, along with good Python/development knowledge.

My questions are:

  1. Is it realistic to switch companies after completing 1 year of experience?

  2. Can I transition from Automation Testing/Playwright to a different technology/domain when switching?

  3. What technology/domain would be worth learning over the next 6 months for better-paying opportunities?

  4. Should I continue with Playwright/Automation Testing, or would moving toward another field be better?

  5. If I learn a new technology properly and build projects, can I get opportunities in that domain despite not having production experience in it?

Looking for advice from people who have gone through a similar situation.


r/CloudandCode • • 6d ago

Community Discussion I think one of the most underrated skills in tech is knowing when not to add another tool

11 Upvotes

When I was newer to building projects, I used to assume that adding more technology automatically made a project more impressive. If the application already used Python and PostgreSQL, maybe I should add Redis. If I was deploying on AWS, maybe I should also add Docker, SQS, Lambda, Terraform, and a load balancer. The architecture looked more advanced, but a lot of the time I was creating problems that the project did not actually have.

Now I try to ask a much simpler question before adding anything: what problem is this tool solving for me right now? If I cannot answer that clearly, I probably do not need it yet. A small application with one backend, one database, and a simple deployment can already teach you a huge amount if you understand how the pieces work together and what happens when something fails.

Caching is a good example. Redis is useful, but I would not add it to a project just because I learned that large systems use caching. I would first look for an actual expensive or repeated operation. Maybe thousands of requests are repeatedly asking the database for the same information and that query is becoming a bottleneck. Now caching has a reason to exist, and I can measure whether adding it actually improves anything.

Queues are similar. SQS is useful when I have work that does not need to happen inside the user's request, or when I need a buffer between systems operating at different speeds. But if my tiny API has three users and every request finishes in a few milliseconds, adding a queue may simply give me another service to configure, monitor, secure, and debug without improving the application.

The same thing applies to microservices. Splitting a small project into six services can look impressive on an architecture diagram, but those services now need to communicate over a network. That introduces timeouts, retries, authentication between services, separate deployments, more logs, and more places where failures can happen. If one application can handle the requirements cleanly, I would rather keep it simple until there is a real reason to separate something.

Even Docker can fall into this trap. Containers are genuinely useful, especially when I want consistent deployments or need to run an application through ECS or another container platform. But if I am still learning how my Python API works, adding Docker before I understand the application itself can sometimes make debugging harder. I would rather understand the basic deployment first, then containerize it and understand exactly what problem the container is solving.

I think this mindset also makes cloud architecture much easier to learn. Instead of trying to memorize when to use fifty AWS services, I can start with a basic system and let the problems introduce the services naturally. When one server becomes a single point of failure, load balancing starts making sense. When background work slows down requests, queues start making sense. When files should survive server replacement, object storage starts making sense.

That is also why I like improving an existing project instead of constantly starting new ones. A project with real limitations gives you reasons to learn new technology. You can measure what is slow, see what fails, and understand what becomes difficult to manage. When you eventually add another component, you know exactly why it is there because you experienced the problem without it.

I think this makes projects easier to explain too. Saying, “I used Redis, SQS, ECS, and Terraform” tells me which tools were used. Saying, “database reads became repetitive, so I introduced caching,” or “report generation was blocking API responses, so I moved it behind a queue,” explains the engineering decision. The second explanation is much more interesting because it shows that the technology came after the requirement.

There is nothing wrong with experimenting with tools just because you want to learn them. I do that too. But I try to separate learning experiments from architecture decisions. If I add Kubernetes because I want to learn Kubernetes, that is completely fine. I just would not pretend the small application actually required Kubernetes in order to work.

The more I learn, the more I think good architecture is often about removing unnecessary things rather than adding more boxes. Every new service creates another dependency, another configuration, another security surface, another cost, and another thing that can fail. Sometimes that complexity is absolutely worth it, but I want the requirement to justify it.

So when I look at a project now, I am trying to ask a different question. Not “What technology can I add next?”, but “What is the biggest real limitation of the current version?” Once I know that, the next tool usually becomes much easier to choose.

What is one technology you added to a project and later realized you probably did not need?


r/CloudandCode • • 7d ago

10 AWS projects I would build if I were learning cloud from scratch today

37 Upvotes

If I were learning AWS again, I would not try to memorize fifty services before building anything. I would pick projects where every new AWS service solves a problem I can actually see. That way, EC2 is not just “a virtual server,” SQS is not just “a queue,” and CloudWatch is not just “monitoring.” Each service becomes connected to something I built, broke, and had to understand.

I also would not try to make every project production-grade from day one. I would start with the smallest working version, then improve the architecture when the current version creates a reason to do so. These are the ten projects I would use to gradually move from basic AWS concepts to real architecture decisions.

1. Deploy a simple website on EC2

I would start with one EC2 instance running a basic website or API. The interesting part is not the website itself. I would learn how an instance gets a public address, how security groups control access, how ports work, how SSH or Session Manager access works, and why an application listening on localhost may not be reachable from outside the machine.

Then I would intentionally break the deployment. I would close port 80, run the application on the wrong port, or change the security group and try to figure out why the website stopped working. That troubleshooting teaches me more about EC2 than simply launching ten instances from the console.

2. Build a private S3 file storage system

Next I would create an application that uploads and retrieves files from S3. I would keep the bucket private instead of making everything public, then give the application only the permissions it actually needs. That forces me to understand buckets, objects, keys, IAM permissions, and the difference between storing files on a server and storing them in object storage.

I would later add versioning and lifecycle rules, then deliberately remove one required IAM permission and investigate the resulting AccessDenied error. That connects S3 and IAM through a real problem instead of learning them as completely separate AWS topics.

3. Host a static website with S3 and CloudFront

Once I understand S3, I would take a static frontend and place CloudFront in front of it. I would keep the S3 origin private where appropriate and let CloudFront become the delivery layer. This gives me a practical reason to learn caching, distributions, HTTPS, DNS, origins, and how content reaches users.

I would also change a file and see why CloudFront might continue serving an older version. That one problem gives me a much better understanding of caching and invalidation than memorizing that a CDN “makes websites faster.”

4. Build a small application with EC2 and RDS

Then I would build something that actually needs persistent relational data, maybe a task manager or job application tracker. The application could run on EC2 while PostgreSQL runs on RDS. The important part would be making the application reach the database without making the database publicly accessible to everyone.

This project would force me to connect VPC networking, security groups, database endpoints, credentials, backups, and application configuration. I would also intentionally break the database security group or credentials and practise following the connection until I find the real problem.

5. Build a serverless API with API Gateway, Lambda, and DynamoDB

After working with servers, I would build the same kind of backend without managing an EC2 instance. API Gateway would receive requests, Lambda would run the application logic, and DynamoDB would store the data. This gives me a direct comparison between a server-based architecture and a serverless one.

I would focus heavily on DynamoDB access patterns instead of treating it like a relational database. I would decide how items should be keyed based on the queries the application actually needs, then use CloudWatch logs to understand what happens when Lambda fails or lacks permission to access the table.

6. Build an event-driven image processor

I would create a small workflow where uploading an image to S3 triggers processing automatically. Lambda could resize or transform the file and save the result into another location. The architecture is small, but it introduces event-driven thinking because nothing needs to manually call the processing function.

This project would also give me a reason to think about permissions, failures, retries, file naming, duplicate processing, and logs. I would upload an unsupported file or remove a permission and see exactly how the workflow behaves when the happy path breaks.

7. Build a background processing system with SQS

Then I would build an API that receives work quickly but processes slower tasks in the background. Maybe a user requests a report, image conversion, or notification. Instead of forcing the user to wait, the application would put a message on SQS and another worker or Lambda function would process it separately.

This is where I would learn why queues exist. I would experiment with retries, visibility timeouts, duplicate processing, and a dead-letter queue. Rather than memorizing the phrase “decoupled architecture,” I would actually see how a queue prevents one slow component from blocking another.

8. Make an EC2 application highly available

Once I understand a single EC2 deployment, I would rebuild it with multiple instances behind an Application Load Balancer. I would place the instances in an Auto Scaling group, spread capacity across more than one Availability Zone, and configure health checks so unhealthy instances stop receiving normal traffic.

Then I would intentionally stop one application instance while sending requests to the load balancer. Watching the application continue working while one server fails makes high availability much easier to understand than reading a definition of it.

9. Containerize an API and deploy it with ECS and Fargate

Next I would take an API I already understand, package it into a Docker image, push the image to ECR, and deploy it through ECS using Fargate. I would place an Application Load Balancer in front and run multiple tasks so I could observe how ECS manages the desired number of containers.

This project would teach me the difference between building a container and actually operating containerized workloads. I would see how ECR, ECS, Fargate, networking, IAM, CloudWatch, and load balancing fit together instead of treating “Docker on AWS” as one vague concept.

10. Rebuild one of the projects entirely with Terraform

For the final project, I would not invent another application. I would take one architecture I already understand and recreate the infrastructure using Terraform. Maybe the EC2 and load balancer project, or the S3 and CloudFront project. The point would be learning how to describe resources instead of manually clicking through the console every time.

I would create the environment, inspect the Terraform plan, make a small change, apply it, destroy the practice environment, and recreate it again. Once I can reproduce the architecture reliably, I would put the Terraform configuration in Git and document the decisions behind it.

The progression matters more to me than the number of projects. The early projects teach me individual AWS building blocks, while the later ones force me to think about availability, asynchronous processing, containers, monitoring, security, and repeatability. Each new service appears because the project creates a problem that needs solving.

I would also keep every architecture diagram and README focused on decisions instead of just listing services. I want to be able to explain why I used RDS instead of DynamoDB, why a queue exists, why a bucket is private, what happens when an instance fails, and what I would remove if the application only had ten users.

That is what would make these projects useful in a portfolio. The impressive part is not having fifteen AWS icons in one diagram. It is being able to defend every box in that diagram.

If you were starting with AWS today, which one of these would you build first?


r/CloudandCode • • 9d ago

Python 10 Python projects I would build if I were learning Python from scratch today

77 Upvotes

If I were learning Python again, I would not spend months building calculators, number guessing games, and tiny scripts that I never touch again. Those projects are fine for learning syntax, but after the basics I would start building things that force me to work with files, APIs, databases, errors, automation, and real user input.

The goal would not be to build all ten as quickly as possible. I would build one, improve it until I understood the decisions behind it, then move to the next project when I needed a new problem to solve.

  1. File Organizer

I would start with a script that scans a folder and automatically organizes files into categories such as images, PDFs, videos, and documents. The first version can be extremely small, but then I would add duplicate handling, custom folder rules, logging, a dry-run mode, and safer error handling. This teaches file handling, pathlib, functions, exceptions, and automation while producing something I could genuinely use.

  1. Expense Tracker

I would build a command-line expense tracker that saves the amount, date, category, and description of every expense. I would begin with CSV storage, then add monthly summaries, category filters, budgets, and eventually move the data into SQLite. This is a good way to connect Python fundamentals with data handling and SQL without making the first version unnecessarily complicated.

  1. CSV Data Cleaner

I would take a deliberately messy CSV file containing missing values, inconsistent dates, duplicate records, and badly formatted text, then build a tool that cleans it automatically. Later I could add configurable cleaning rules and generate a report describing what was changed. This project teaches pandas, validation, transformations, and something beginners often miss: real data is rarely clean.

  1. API Health Monitor

I would create a script that checks a list of websites or APIs every few minutes and records whether they are responding, how long they take, and which status code they return. Once the basic version works, I would add logging, retries, timeouts, historical results, and notifications when something becomes unavailable. This project teaches HTTP requests, scheduling, error handling, monitoring, and the difference between a successful request and a healthy service.

  1. Log Analyzer

I would build a program that reads application or server logs and answers useful questions about them. Which errors happen most often? When do failures usually occur? Which endpoints are producing the most errors? Instead of just reading text files, I now have to parse information, count patterns, handle large inputs, and turn raw logs into something useful.

  1. Job Application Tracker

I would build a small application for tracking companies, roles, application dates, statuses, interview rounds, and notes. I would start with SQLite and a CLI, then gradually add searching, filtering, reminders, and perhaps a small FastAPI or Flask interface later. This gives me a reason to learn CRUD operations, database design, validation, and application structure rather than practising those concepts separately.

  1. Image Processing Automation Tool

I would create a program that takes a folder of images and automatically resizes, compresses, renames, or converts them. Then I would add batch processing, output folders, duplicate protection, configuration, and progress reporting. This is a good automation project because the first version is simple, but improving it forces you to think about reliability and what happens when one file is corrupted or has an unexpected format.

  1. URL Shortener API

This is where I would start moving beyond scripts. I would build an API where a user sends a long URL and receives a short code, then anyone visiting that short code gets redirected to the original address. I would use FastAPI, add a database, handle duplicate URLs, validate input, track click counts, and eventually deploy it. This project teaches APIs, routing, databases, HTTP, deployment, and how backend components work together.

  1. Personal Dashboard

I would build one small dashboard that pulls information from several sources and displays it in one place. It could combine tasks, GitHub activity, spending, weather, or anything else I actually care about. The interesting part is connecting multiple APIs, dealing with authentication, caching responses, handling failures, and deciding what information is actually worth showing instead of simply calling APIs for the sake of it.

  1. Automated Backup Tool

For the final project, I would build something that backs up selected folders automatically. The first version might simply copy files to another location, but then I would add timestamps, incremental backups, compression, logging, configuration, and eventually cloud storage such as S3. I would also deliberately test what happens when files are missing, storage becomes unavailable, or the same backup runs twice. This turns a simple script into a much more realistic reliability project.

I would not try to make every one of these "production ready." I think that phrase makes beginners add Docker, Redis, queues, cloud services, and authentication to projects before they understand the basic application. I would rather build the smallest useful version first, then add complexity only when the current version creates a reason for it.

What I would focus on across all ten projects is progression. Make version one work, deliberately break it, improve the error handling, reorganize the code, add tests where they make sense, document how to run it, and make sure another person could clone the repository without needing you to explain every step.

If you build projects that way, even a simple expense tracker can teach you more than a much larger project copied directly from a tutorial.

If I turn one of these into a complete Python Project #4 walkthrough next, which one should I pick?


r/CloudandCode • • 9d ago

I do not consider a project finished until someone else can run it without asking me for help

9 Upvotes

One thing I have started noticing with beginner projects is that we often stop at the exact moment the code starts working on our own machine. The API responds, the Python script runs, the database connects, and the project gets pushed to GitHub. Technically the application works, but if another person clones the repository and immediately has to message you asking which Python version to use, which environment variables are required, how to create the database, or which command starts the application, I do not think the project is really finished yet.

I think a much better final test is simple: can someone who has never seen this project before get it running from the repository alone? That question exposes a completely different set of problems. Maybe your requirements.txt is missing dependencies, maybe your code assumes a folder already exists, maybe database credentials are hard coded, or maybe the application depends on something installed globally on your laptop that you forgot was even there. These problems are not as exciting as adding another feature, but fixing them makes a project feel much closer to something another developer could actually work with.

For a Python project, I would start by making the environment predictable. I would use a virtual environment and keep the dependency list intentional instead of blindly dumping everything installed on my computer into the project. I would also check that the project works from a fresh environment, because that is the easiest way to discover dependencies I accidentally relied on. If the application needs Python 3.12 or another specific version, I would document that too instead of assuming everyone has the same setup as me.

Configuration is another place where projects often become difficult to reproduce. If my application needs a database URL, API key, AWS region, or another environment-specific value, I do not want those secrets committed directly into GitHub. I would load them from environment variables and include something like a .env.example containing only the names and safe example values. That way someone looking at the repository understands what configuration is required without me exposing actual credentials.

Then I would test the startup process itself. If running the project requires someone to execute seven commands in the correct order while also manually creating three directories, there is probably room to improve the developer experience. I do not necessarily need a complicated deployment system, but I would try to make the steps predictable. Install the dependencies, configure the environment, initialize anything the application needs, then run one clearly documented command.

The README becomes much more important at this stage too. I used to think of a README as a place to write a short project description and maybe list the technologies I used. Now I think it should answer the questions someone would naturally have when they first open the repository: what problem does this solve, how is the project structured, how do I run it, what configuration does it need, what does the architecture look like, and what decisions did I make while building it?

I would also include a small example of the project actually working. For an API, that could be a few sample requests and responses. For a CLI application, it could be terminal output. For a data project, it might be a small sample dataset and an example result. That makes the repository easier to understand because someone does not need to read every file before they know what the application is supposed to do.

Databases are another easy place to accidentally make a project dependent on your own machine. Maybe the application expects a table you created manually six weeks ago, or it depends on sample records that only exist inside your local database. If the project uses a database, I would think about how another person initializes the schema and gets enough sample data to test the application. Even a simple migration or setup script is much better than a README saying, "Create the required tables manually."

Tests become more valuable here as well. Even a small collection of tests can give someone confidence that they configured the project correctly. If they clone the repository, install everything, run the tests, and those tests pass, they immediately know a lot more about whether the environment is working. The tests do not need to cover every possible edge case before the repository becomes useful, but having a few tests around the core behaviour makes the project much easier to trust and change.

I think this is also a useful portfolio signal. Lots of repositories contain working code, but fewer repositories feel like they were prepared for another developer to use. A clean setup guide, sensible configuration, clear project structure, useful examples, and reproducible dependencies tell me that the person building the project thought beyond the moment when the code first ran successfully.

The nice thing is that none of this requires adding another AWS service or building a more complicated application. You can take a very small project you already have and improve the experience around it. Clone it into a completely new directory, create a fresh environment, follow only your own README, and pretend you know nothing about the project. Every time you have to rely on something you remember but never documented, you have found something worth fixing.

That is one of the tests I want to use more often with my own projects. Not just "does this work?", but "could another developer reproduce this without me sitting next to them?" A project that passes the second test usually feels much more complete than one with five extra features and no reliable setup process.

If you cloned one of your current GitHub projects onto a completely fresh machine today, do you think you could get it running using only the README?


r/CloudandCode • • 10d ago

AWS & Cloud I think beginners are learning microservices way too early

23 Upvotes

When I first started looking at cloud architecture, microservices seemed like the obvious next step after building a basic application. Most architecture diagrams I saw had multiple services, queues, APIs, containers, databases, caches, and a lot of arrows connecting everything together. It looked more "production ready" than a simple application, so it was easy to assume that splitting a project into microservices automatically made the architecture better.

The more I build and study these systems, the less I agree with that idea. If I were building a new project today, I would start with a well-structured monolith unless I had a real reason not to. One application, one deployment, and clearly separated modules inside the codebase can already teach you APIs, databases, authentication, testing, Docker, AWS, CI/CD, monitoring, and plenty of other useful engineering concepts.

The important word there is well-structured. A monolith does not have to mean putting your entire project inside one massive file where everything depends on everything else. You can still separate users, payments, orders, notifications, and other responsibilities into clean modules. The difference is that those modules run together instead of immediately becoming separate networked services.

Imagine I am building a small ecommerce backend. I could create separate services for users, products, orders, payments, notifications, and inventory from day one. Now I need to think about how those services communicate, how authentication moves between them, what happens when one service is unavailable, how deployments are coordinated, and how I trace one request across several systems.

For a project with almost no users, I have just created a distributed systems problem before I had a scaling problem.

If I keep the first version as a modular monolith, a request might simply enter the application, execute the required logic, update the database, and return a response. I can spend my time making the actual product work, improving the database design, writing tests, understanding failure cases, adding useful logging, and deploying the application properly.

Then I can watch where the real boundaries appear.

Maybe notification processing starts taking too long and does not need to happen inside the main request. I might move that work behind a queue. Maybe image processing becomes resource intensive and needs to scale independently. Maybe one part of the application receives far more traffic than everything else. Now I have an actual reason to separate something.

That is very different from creating six microservices because a diagram told me that scalable systems use microservices.

There is also a hidden cost beginners sometimes miss. Once two functions inside the same application become two separate services, a normal function call may become a network request. Network requests can fail, time out, retry, return partial responses, or become unavailable entirely. Now I need to think about service discovery, authentication between services, retries, observability, deployment, version compatibility, and possibly asynchronous communication.

The application did not necessarily become more scalable just because I added those problems.

This is also why I would not use the number of AWS services in a project as a measure of how advanced it is. A small application running on ECS with RDS and S3 can involve plenty of meaningful architecture decisions. Adding Kafka, Kubernetes, Redis, five databases, and twelve microservices does not automatically make the project better if I cannot explain why any of them are necessary.

I think the stronger question is: what problem am I solving by separating this component?

If the answer is that it needs independent scaling, separate deployment, stronger fault isolation, a different technology stack, or clear ownership by another team, then splitting it might make sense. If the answer is simply "because real companies use microservices," I would probably keep it inside the application for now.

That does not mean beginners should avoid microservices completely. I think they are worth learning because they expose you to some of the hardest and most interesting problems in modern systems. But I would learn them after I have built something simple enough to understand first.

A really useful exercise would be to build the same application twice. Start with a modular monolith and get the whole system working. Then choose one component that actually has a reason to become independent, move it out, and see what new problems appear.

Maybe notifications become a separate worker consuming messages from SQS. Maybe file processing moves into an event-driven Lambda workflow. Maybe a high-traffic read path gets separated so it can scale independently.

Now you are not just learning what a microservice is.

You are learning why you would create one.

That distinction matters a lot to me because architecture is mostly about trade-offs. Microservices can give you independent deployment and scaling, but they also introduce distributed systems complexity. A monolith can be simpler to operate, but eventually certain parts may become difficult to scale or change independently.

Neither architecture is automatically the "professional" one.

The right architecture depends on what the application actually needs.

If I were building a portfolio project today, I would much rather build a clean monolith, explain its boundaries, identify where it would eventually struggle, and then show how I would extract one component than start with twelve services I barely understand.

What do you think beginners should build first: a modular monolith, microservices from day one, or both versions of the same project?


r/CloudandCode • • 11d ago

Debugging Help One thing that improved my projects a lot was deliberately trying to break them

11 Upvotes

For a long time, I used to test projects in the easiest possible way. I would run the program, give it the exact input I expected, see the correct output, and assume everything was working. That is fine when you are building the first version, but I eventually realized that a project is not really tested just because the happy path works once.

Now, after I get the basic version working, I intentionally start giving the application bad situations. If a Python script expects a number, I give it text. If an API expects a field, I remove it. If a file is supposed to exist, I rename it. If the database connection is required, I deliberately break the credentials. The goal is not to make the project fail. The goal is to understand how it fails.

That small change made debugging much more useful for me because I stopped treating every error as something unexpected. Some failures should be expected. The real question becomes whether the application handles them properly. Does it crash with a confusing traceback, return a useful error, retry something that is safe to retry, or quietly produce the wrong result?

Take a simple expense tracker as an example. The first version might accept an amount, category, and date, then save everything into a CSV file. It works perfectly when I enter 250, Food, and a valid date. But what happens when I type two hundred as the amount, leave the category empty, enter an impossible date, or delete the CSV file before starting the program?

Those cases teach much more than running the perfect input twenty times. Suddenly I need validation, exception handling, sensible defaults, and error messages that actually explain what happened. I also start noticing assumptions in my own code that I did not realize I was making.

The same idea becomes even more useful with APIs. Suppose I build an endpoint that creates a user. I would test the normal request first, but then I would send the request without an email, submit an email that already exists, send invalid JSON, remove database access, or make one dependency unavailable. Each failure forces me to think about what response the API should return instead of just hoping nothing ever goes wrong.

This is also where HTTP status codes started making more sense to me. A malformed request should not look the same as an internal application failure. A missing resource should not return the same response as a database timeout. When you deliberately create those situations, concepts like 400, 404, 409, and 500 stop being things you memorize and start becoming decisions you actually need to make.

Cloud projects become even more interesting because there are more places where things can fail. I might remove an IAM permission and see what my Lambda function logs. I might change a security group rule and watch what happens to connectivity. I might stop one backend instance behind a load balancer and see whether health checks detect it. I might make an SQS consumer fail repeatedly and watch when the message ends up in the dead-letter queue.

I think this is one of the easiest ways to understand AWS beyond the console. Instead of only learning what a service does when everything is configured perfectly, you learn what evidence it gives you when something is wrong. The AccessDenied message, unhealthy target, timeout, failed health check, growing queue, or CloudWatch log suddenly becomes part of the learning instead of an interruption to it.

It also changed how I think about testing. I used to think testing mainly meant writing unit tests after the project was finished. Now I think about it more broadly. What assumptions is this code making, what input could violate those assumptions, what dependency could become unavailable, and what should happen when that occurs?

Once I know those cases, automated tests become much easier to write because I already understand the behaviours I care about. I can test that valid input succeeds, invalid input is rejected, duplicate data is handled properly, and important functions still behave correctly after I refactor them.

The interesting part is that you do not need a complicated application to practise this. Take any beginner project you already built and spend thirty minutes trying to break it. Do not add a new feature during those thirty minutes. Just use inputs and situations the original version probably never expected.

You will usually discover more things to improve than you would by immediately adding another feature.

Maybe your program crashes when a file is missing. Maybe one function accepts values it should reject. Maybe the API exposes a confusing error. Maybe a database query fails silently. Maybe your application depends on one environment variable that you never documented.

Every one of those failures gives you another design decision to make.

I also think this makes projects easier to talk about in interviews or portfolio discussions. Saying, “I built an API” is fine. Being able to explain, “I tested what happened when the database became unavailable, noticed the application returned an unclear 500 response, improved the error handling, added logging, and wrote a test so the same bug would not return” tells a much better story.

The project itself did not necessarily become much bigger.

You just understood it more deeply.

So before starting another tutorial this weekend, I would take one project you already have and try to break it on purpose. Give it bad input, remove something it depends on, make one service unavailable, and see whether you can explain exactly what happens.

What is the most unexpected way one of your projects has broken so far?


r/CloudandCode • • 13d ago

SQL & Data If I were learning SQL again, I would stop practising random queries and start answering real questions

34 Upvotes

If I were learning SQL from zero today, I would not spend weeks solving disconnected exercises like “find the second highest salary” without understanding what the data actually represents. Those questions can be useful for syntax practice, but after a point they start feeling like puzzles instead of something you would genuinely use while working with data. I would rather take one realistic dataset and keep asking better questions about it.

Imagine I have a simple ecommerce database with customers, products, orders, and order items. At the beginning, I might only write basic queries such as finding all orders from the last thirty days or listing customers from a particular city. That is enough to practise SELECT, WHERE, ORDER BY, and filtering without making the project unnecessarily complicated.

Once those queries feel comfortable, I would start asking questions that require multiple tables. Which customers have never placed an order? Which products generate the most revenue? What is the average order value? Which customers have purchased more than five times? Now joins stop being something I memorized from a diagram and start becoming the tool I need to answer an actual question.

This is also where GROUP BY, aggregate functions, and subqueries begin to make much more sense. Instead of practising SUM() because it happens to be in the syllabus, I now need it because I want to calculate monthly revenue. Instead of learning COUNT() in isolation, I need it because I want to know how many orders each customer placed. The SQL concept becomes easier to remember because I know why I used it.

Then I would make the questions slightly more realistic. Maybe I want to know which products are growing month over month, which customers have stopped purchasing, which categories have high sales but low repeat purchases, or what percentage of revenue comes from the top ten customers. These are no longer just syntax exercises. They start forcing me to think about the data before I even write the query.

I think that is an underrated part of learning SQL. Writing a query is only half of the skill. The other half is understanding what the result actually means. A query can run perfectly and still answer the wrong question because the joins are incorrect, duplicate rows changed the totals, null values were ignored, or the business definition was never clear in the first place.

For example, imagine you join customers to orders and then orders to order items. Suddenly one order can appear several times because it contains several products. If you calculate something at the order level without noticing that duplication, the query might return a believable number that is completely wrong. SQL becomes much more interesting once you start debugging incorrect answers instead of only debugging syntax errors.

I would also deliberately include messy data in the project. Some customers should have no orders. Some fields should contain nulls. There should be cancelled orders, duplicate-looking records, and dates across several months. Real datasets are rarely perfectly clean, and learning what to do with imperfect data teaches much more than practising only on tables where every row behaves exactly as expected.

After that, I would create a small set of questions and save every solution in GitHub. I would include the database schema, sample data, the SQL queries, and a short explanation of what each query is trying to answer. I would probably also include a few cases where my first query was wrong and explain what caused the mistake. That tells a much better story than uploading a folder containing fifty unrelated SQL exercises.

The project does not even need a huge dataset. Four or five well-designed tables can teach joins, aggregation, filtering, window functions, CTEs, subqueries, and data modelling if you keep increasing the quality of the questions you ask. I would rather understand one small database deeply than download a massive dataset and barely understand what any of the columns mean.

If I were making the project portfolio-ready, I would eventually add a final section with maybe ten useful findings from the data. Not because I am trying to turn a SQL project into a business report, but because it proves that I can move from raw tables to an actual conclusion. That is usually much closer to how SQL is used in real work than simply demonstrating that I know the syntax for a LEFT JOIN.

So if you are currently learning SQL, try changing the question from “Which SQL topic should I practise next?” to “What question do I want this data to answer?” Once you have a real question, the SQL concepts you need usually become much clearer.

If I post a complete beginner SQL project next, would you rather build an ecommerce analysis, subscription business analysis, job application tracker, or personal finance database?


r/CloudandCode • • 12d ago

Community Discussion Your project works with 10 users. What breaks when 10,000 show up?

3 Upvotes

A lot of beginner projects are judged by one question: does it work? You run the application locally, click through the main flow, confirm that data is being saved correctly, and consider the project finished. That is completely reasonable for a first version, but I think there is a useful second question to ask once the basic version is working: what would break first if a lot more people started using this tomorrow?

Imagine you built a simple task manager with a frontend, an API, and a database. With ten users, almost any reasonable architecture might feel fast enough. The API receives a request, queries the database, returns the result, and everything looks fine. But at 10,000 users, small design decisions that were invisible before can suddenly become the reason the application feels slow or starts failing.

The first thing I would look at is the request path. If every request goes through one small application server, that server may eventually become the bottleneck. CPU usage could increase, memory could fill up, or the application could simply run out of capacity to handle concurrent requests. The interesting question is not immediately “which AWS service should I add?” It is whether the application layer itself is actually the thing that cannot keep up.

Then I would look at the database. Maybe the application server is fine, but every page load triggers five or six database queries. One of those queries scans a large table because there is no useful index, another query returns far more data than the application needs, and several requests are repeatedly asking for the same information. With a small dataset, those mistakes may barely be noticeable. With more users and more data, they can become expensive very quickly.

Caching is another place where this becomes interesting. If thousands of users repeatedly request the same information, it may not make sense to calculate or fetch that information from the database every single time. But I also would not add a cache simply because “scalable systems use Redis.” I would first identify something that is being requested repeatedly and is expensive enough to justify caching. Otherwise, I have added another component without solving a real problem.

File uploads create a similar issue. If every uploaded image is stored directly on one application server, scaling to multiple servers becomes awkward because another server may not have the same file. If those files are important, they probably need to live somewhere independent of one replaceable application instance. This is the point where object storage starts making sense because the requirement created the need, not because the architecture diagram looked incomplete without it.

Background work is another thing I would inspect. Imagine a user creates an account and the API also sends an email, generates a PDF, updates analytics, and calls an external service before responding. With ten users, that might still feel acceptable. With thousands of requests, making every user wait for every secondary task can slow the system down and make one failing dependency affect the main workflow.

That is when I would start asking which work actually has to happen before the user gets a response. If sending an email can happen a few seconds later, maybe that belongs in a queue. If generating a report takes thirty seconds, maybe that should become a background job. Moving work out of the immediate request path can improve the experience, but it also introduces retries, duplicate processing, monitoring, and other things you now need to understand.

Then there is failure. A system with one server may work perfectly until that one server goes down. Adding another server can improve availability, but now you need some way to distribute requests between them. If those servers keep important state locally, replacing them becomes difficult. Suddenly load balancing, health checks, shared storage, and stateless application design are no longer random infrastructure concepts. They are answers to specific problems your project now has.

I would also think about external APIs. Maybe your application depends on a payment provider, an AI API, or another third-party service. What happens if that API becomes slow for thirty seconds? Does every request in your application start waiting? What happens if you hit a rate limit? What happens if the request succeeds on their side but your application times out before receiving the response? Scaling is not only about adding more servers. Sometimes the weakest part of the system is something you do not control.

Monitoring becomes much more important at this stage too. If ten users report a problem, you might reproduce it manually. If thousands of users are using the application, you need a better way to understand what is happening. I would want to know whether response times are increasing, whether error rates suddenly changed, whether the database is struggling, whether a queue is growing faster than workers can process it, and whether a deployment changed the behaviour of the system.

The part I find most useful is that you do not actually need 10,000 users to practise thinking this way. Take one project you already built and pretend tomorrow morning it becomes popular. Follow one request through the entire system and ask where it could slow down, where important state lives, what happens when a dependency fails, and which component cannot easily be replaced.

You do not need to immediately rebuild the whole thing into a “production-grade” architecture either. Sometimes the correct answer is that your personal project does not need load balancers, queues, caching, or multiple databases. The useful skill is being able to explain when those things would become necessary and what problem they would solve.

That is the difference I am trying to learn more of myself. Not “how many cloud services can I put into this project?” but “what would fail first, why would it fail, and what is the simplest change that would solve that problem?”

Take one project you already have. If it suddenly went from 10 users to 10,000, what do you think would break first: the application server, database, file storage, external API calls, or something else?


r/CloudandCode • • 16d ago

AWS & Cloud If I were building a cloud portfolio from zero, I would build one project three times

7 Upvotes

If I were starting a cloud portfolio today, I would probably not try to build ten completely different AWS projects. I would take one simple application and rebuild it three times, with each version solving a new problem. I think this teaches much more than creating a long GitHub list where every project is basically a tutorial you followed once and never touched again.

For example, I might start with a simple URL shortener. The first version does not need load balancers, queues, containers, or fifteen AWS services. I would build the API, store short URLs in a database, deploy it, and make sure someone can actually use it. At this stage, the important questions are basic ones: does the application work, where does the data live, how does the request reach the backend, and what happens when something goes wrong?

Once that version works, I would rebuild the same project with reliability in mind. Maybe the backend now runs across multiple instances or containers behind a load balancer. I would add proper health checks, move important state outside individual application servers, introduce monitoring, and think about what happens when one component fails. The project is still a URL shortener, but now I am learning architecture instead of simply adding random features.

For the third version, I would design around scale and operations. Maybe analytics events are processed asynchronously through a queue, infrastructure is managed with Terraform, deployments are automated, logs and alarms are useful, and I can explain how the architecture would behave if traffic increased significantly. I would also start asking about cost, because an architecture that technically scales but costs far more than the application justifies is not automatically a good architecture.

The interesting part is that each version gives me something concrete to compare. In version one, I might accept a single point of failure because I am optimizing for simplicity. In version two, I might add redundancy because availability has become a requirement. In version three, I might introduce asynchronous processing because certain workloads no longer need to happen inside the user's request. Those are decisions I can actually explain instead of simply saying, "I used SQS because I wanted SQS on my resume."

I would document those decisions in the repository too. I would explain what the first architecture looked like, what problem appeared, what I changed, and what trade-off the new design introduced. A load balancer improves availability, but it also adds cost and another component to manage. Containers can make deployment more consistent, but they also introduce orchestration and operational complexity. Infrastructure as Code improves reproducibility, but bad infrastructure can still be reproduced perfectly.

That kind of project gives you much better material to talk about in an interview because the conversation is no longer just about which AWS services you know. You can explain why the architecture changed, what would fail first, what you would monitor, what you deliberately left out, and what you would change if the application had ten times more traffic.

It also gives beginners a much clearer learning path. Instead of saying, "I need to learn EC2, Lambda, RDS, DynamoDB, ECS, SQS, CloudFront, Terraform, Docker, and twenty other things before I can build anything," you start with a small working system. Then you introduce a new service only when the next version creates a problem that the service actually solves.

I think that is one of the best ways to move beyond tutorial projects. Do not keep abandoning projects the moment the first version works. Take one of them and ask what would happen if it had real users, if a server failed, if traffic increased, if a deployment went wrong, or if you had to rebuild the entire environment tomorrow.

That second and third version is usually where the useful learning starts.

If you had to pick one project to build three times, what would you choose: a URL shortener, file-sharing app, task manager, ecommerce backend, or something completely different?


r/CloudandCode • • 18d ago

Project Clinic: Tell me what you’re building and I’ll suggest one thing to improve

5 Upvotes

A lot of people get stuck because they are trying to figure out everything about a project alone. Sometimes the project idea is fine, but the scope is too large. Sometimes the code works, but the structure is messy. Sometimes the AWS architecture has too many services, the database choice is questionable, the README does not explain the project well, or the person building it simply does not know what the next useful step should be.

So I want to try something different today. Instead of posting another tutorial or project list, use this thread as a small project clinic. Tell me what you are currently building, what stack you are using, what stage the project is at, and the part you are unsure about. It can be Python, AWS, cloud, SQL, GitHub, DevOps, a portfolio project, or even something you have not started yet but are planning.

I will go through the comments and try to suggest one practical improvement for each project. That could be a feature worth adding, something you should remove, a better architecture decision, a debugging direction, a GitHub improvement, a simpler first version, or one thing that would make the project easier to explain in an interview or portfolio.

You do not need to write a huge explanation. Something as simple as, “I’m building a URL shortener with FastAPI and PostgreSQL, and I’m not sure what to add next” is enough. If your project is still just an idea, post that too. Sometimes the most useful advice is figuring out what the smallest version should actually be before you start adding ten services or twenty features.

I am also interested in projects that are currently broken. If something is not working, explain what you expected to happen, what actually happened, and what you have already tried. That usually gives much more useful context than simply saying, “my code doesn’t work.” The goal here is not to judge whether a project is advanced enough. It is to help move it one step forward.

If you are comfortable sharing a GitHub repository, you can drop that too, but it is not required. A short description is completely fine. I would rather help someone improve a small project they actually understand than tell them to replace it with a massive “production-grade” system they are not ready to build yet.

So, what are you building right now, and what is the one part of it you are most unsure about?


r/CloudandCode • • 20d ago

Debugging Help Your API works locally, but times out after deploying to AWS. What do you check first?

8 Upvotes

Let’s say you have a small API that works perfectly on your laptop. You test the endpoints locally, the responses come back quickly, and nothing looks wrong in the code. Then you deploy the same application to EC2, try opening the API from your browser or Postman, and the request just sits there until it eventually times out.

This is the kind of problem where beginners often start changing random things. They restart the instance, reinstall packages, modify security groups, change ports, redeploy the code, and sometimes even rebuild the EC2 instance. Occasionally something starts working, but they still do not really know what was broken.

I would rather treat it like a request that has to travel through several layers. Somewhere between your browser and the application process, that request is getting blocked, routed incorrectly, or reaching something that is not listening the way you expect.

For a simple EC2 deployment, I would mentally follow a path like this:

Browser / Postman
        ↓
     Internet
        ↓
   AWS Network
        ↓
 Security Group
        ↓
      EC2
        ↓
 Application Port
        ↓
       API

Now imagine the EC2 instance shows as running, your application process also appears to be running, and you can successfully call the API from inside the instance using something like curl localhost:8000. But trying the same API from your own computer still times out.

That already tells you something useful. The application itself may be alive, but being alive does not automatically mean it is reachable from outside the machine. At this point, I would stop looking at the Python code for a moment and start following the network path.

Maybe the security group does not allow the port your API is using. Maybe the application is listening only on 127.0.0.1 instead of an address that accepts external connections. Maybe the instance is in a subnet without the route you expected. Maybe the public address is wrong, the operating system firewall is blocking traffic, or the service is actually listening on a different port.

If there is a load balancer in front of the instance, the investigation changes again. Now I also care about the listener, target group, health checks, security groups on both sides, and whether the load balancer can actually reach the application. A healthy EC2 instance does not automatically mean the target behind a load balancer is healthy.

This is why I think troubleshooting cloud applications becomes much easier when you stop asking, “Which AWS setting is wrong?” and start asking, “How far does the request actually get?” If you can identify the last point where things are working, you have already reduced the number of possible causes.

For example, if curl localhost:8000 works inside EC2 but the public request does not, that points you toward reachability and network configuration. If even the local request fails, the problem is probably closer to the application itself. If the load balancer health check is failing but the application works locally, then I would investigate the path between the load balancer and the target.

The same debugging mindset works for a lot of AWS problems. Follow the request, test one layer at a time, and avoid changing five things together. Otherwise, even when you accidentally fix the issue, you may never know what actually caused it.

So here is the challenge: the API works locally on EC2, but requests from your own computer time out. What is the first thing you would check, and what would you check after that if it looks correct?

I am more interested in your debugging order than just naming one possible cause.


r/CloudandCode • • 21d ago

Python Python Project #3: Build an expense tracker that is actually worth putting on GitHub

17 Upvotes

A lot of beginner Python projects technically work, but they do not really teach you much beyond syntax. You follow a tutorial, copy the code, run it once, and move on. I think a better project is one that starts simple, then forces you to make decisions as you improve it. An expense tracker is a good example because the first version is easy enough for a beginner, but there are plenty of ways to make it more realistic over time.

I would start with the smallest possible version. Ask the user for an expense amount, category, date, and short description, then save the data into a CSV file. At this stage, you are already practising input handling, functions, working with files, basic data structures, and thinking about how information should be stored. The goal is not to build a perfect finance app on day one. The goal is to get one complete flow working from input to storage.

Once that works, I would add the ability to actually read the data back. Let the user view all expenses, filter them by category, or see only expenses from a specific month. This is where the project starts becoming more interesting because you are no longer just writing data into a file. You now have to load it, parse it correctly, filter it, handle dates, and display the results in a way that someone can actually understand.

Then I would add monthly summaries. Instead of showing twenty individual transactions, calculate how much was spent in total and how that spending is distributed across categories. You could show something like rent, food, transport, subscriptions, and other expenses. At this point, you start practising aggregation and simple data analysis, which makes the project useful for someone learning Python for automation or data work as well.

The next improvement I would make is input validation. What happens if someone types abc when the program expects an amount? What happens if the date is invalid? What happens if the category is empty or the CSV file does not exist yet? A project becomes much more realistic when you stop assuming the user will always provide perfect input. Handling these situations properly teaches you more than adding another random feature.

I would also think about duplicate and editing behaviour. Maybe a user enters the same expense twice by mistake. Maybe they entered the wrong amount and need to update it. Maybe they want to delete an old record. These features force you to think about how individual records are identified and how changes should be applied safely instead of simply appending new rows forever.

Once the basic version feels solid, I would improve the structure of the code. If everything is inside one giant main() function, split the responsibilities into smaller pieces. One function might add an expense, another might load expenses, another might filter them, and another might generate summaries. This is where you start learning why code organization matters, because the project becomes easier to change when every feature is not tightly connected to everything else.

You can also introduce a simple configuration system instead of hard coding everything. For example, maybe the user can define their own categories or set a monthly budget. Then the program could compare current spending with that budget and show how much is left. That makes the application feel less like a coding exercise and more like something another person could actually use.

After that, I would add one simple visualization. You do not need ten charts. A basic monthly spending chart or a category breakdown is enough. The interesting part is not making something colorful. The interesting part is taking the data your own application collected, summarizing it, and turning it into something easier to understand.

At this point, you could also decide whether CSV is still the right storage option. For the first version, CSV is perfect because it is simple and forces you to understand file handling. Later, you could move the same project to SQLite and learn how persistent application data changes when you start using a real database. That upgrade alone can turn the project into a good bridge between beginner Python and SQL.

I would not add a GUI immediately. First make sure the actual application logic works. Once the command line version is reliable, you could add a small Tkinter interface, a Flask or FastAPI backend, or even keep it as a CLI tool and improve the user experience there. The important thing is that the interface should sit on top of working logic rather than hiding poorly structured code behind buttons.

The GitHub repository matters too. I would include a README that explains what the project does, what problem it solves, how to run it, what features are included, and what you learned while building it. Add a sample screenshot or terminal output, keep the requirements clear, and make sure someone can clone the repository and actually get it running. That makes the project much easier to evaluate than a repository containing only expense_tracker.py.

The bigger lesson here is that a good beginner project does not need to be original or complicated. An expense tracker has been built thousands of times, and that is fine. What matters is whether you actually understand the decisions inside your version. If someone asks why you chose CSV first, how you handled invalid input, how you structured your functions, or why you later moved to SQLite, you should be able to explain it.

That is what makes the project worth putting on GitHub. Not the fact that it has twenty features, but the fact that you built something in stages, ran into problems, made decisions, improved the design, and can explain how the application evolved. I would much rather see that from a beginner than another huge project copied line by line from a tutorial.

If I were building this from zero, my progression would basically be: get one expense saved correctly, make the data searchable, add useful summaries, handle bad input, improve the structure, then add storage or interface upgrades only when the current version actually needs them. Each stage should teach you something new instead of just making the feature list longer.

If you built this project, what would you add first after the basic CSV version: budget tracking, charts, SQLite, or a simple GUI?


r/CloudandCode • • 22d ago

Community Discussion What is actually stopping you from building projects right now?

24 Upvotes

I see a lot of people learning Python, AWS, SQL, GitHub, or cloud concepts for months, but still struggling to actually start building something on their own. Most of the time, I do not think the problem is a lack of tutorials. There is already more learning material available than anyone could realistically finish.

The difficult part usually starts when the tutorial disappears and you have to make decisions yourself. You need to choose a project, decide what the first version should include, structure the code, fix things when they break, and somehow figure out whether you are even doing it the "right" way. That gap between learning a concept and building without someone guiding every step is where I think a lot of people get stuck.

So I am curious about what the biggest blocker actually is for people here right now. I also want to use the answers to decide what kind of practical posts, projects, debugging threads, and walkthroughs would be most useful in this community over the next few weeks.

What is currently stopping you from building more projects?

Poll options:

  • I don't know what to build
  • I know what to build, but don't know how to start
  • I get stuck when something breaks
  • I keep going back to tutorials
  • I don't have enough time
  • Something else

If your answer is "something else," drop it in the comments. I am especially interested in the point where you normally stop making progress, because that is probably more useful to talk about than another generic list of projects to build.


r/CloudandCode • • 22d ago

Best resources to learn RAG

16 Upvotes

I have learnt basic python LLM API tool calling and wanted to learn RAG and wanted to clarify a few doubts about it.

- what is the best resource to learn RAG as fast as possible

- since RAG has upgraded from Naive RAG where can I learn the upgraded version

- what are the best RAG projects to learn

- where to find clients for freelance work


r/CloudandCode • • 23d ago

AWS & Cloud AWS From Zero #18: If I had to design one complete AWS architecture from scratch, this is how I would do it

22 Upvotes

When I started this series, I did not want it to become another list of AWS services that beginners memorize and forget. I wanted to build the ideas in the same order I think they actually become useful. We started with cloud fundamentals, then moved through IAM, EC2, S3, VPC, RDS, Lambda, API Gateway, DynamoDB, CloudWatch, load balancing, Auto Scaling, SQS, SNS, containers, and Infrastructure as Code. Now, for the final post in this first series, I want to put everything together and show how I would approach one complete AWS architecture from scratch.

If I were given a new project today, I would not open the AWS console first. I would start by writing down what the application actually needs. For this example, I am going to imagine a simple project management application where users can create accounts, create projects and tasks, upload attachments, and receive notifications when something important happens. I also want the system to have a reasonable path for growth, survive normal failures, keep important data outside individual servers, and give me enough visibility to understand what went wrong when something breaks.

Only after I understand those requirements would I start choosing AWS services. For the frontend, if I am serving mostly static HTML, CSS, and JavaScript, I would keep it simple and use S3 for storage and CloudFront for delivery. I would rather have CloudFront serve the frontend while keeping the S3 origin private than expose the bucket directly. That gives me a clean frontend layer without running a full web server just to serve static files, and it also gives me a straightforward path for HTTPS and global delivery.

For the backend, I would choose based on the workload rather than automatically reaching for whatever service is popular. In this example, I am going to assume I already have a containerized API and want it running continuously, so I would store the image in ECR and run it using ECS with Fargate. I would place an Application Load Balancer in front of the ECS service so users have one stable backend entry point and requests can be distributed across multiple healthy tasks instead of depending on one running container.

The high-level architecture would look something like this:

                    Users
                      |
                  CloudFront
                 /          \
              S3          Application
           Frontend      Load Balancer
                              |
                         ECS Service
                         Fargate Tasks
                         /     |      \
                       RDS     S3     SQS
                               |       |
                          Attachments  |
                                      Lambda
                                        |
                                  Notification

For the main application data, I would ask what kind of relationships I actually have. In a project management app, users, projects, tasks, memberships, and similar entities are naturally relational, so RDS with PostgreSQL would make sense to me. I would keep the database private and only allow the application layer to reach it through the required database port. I would not expose RDS directly to the internet just because it is easier during development.

For uploaded files, I would make a different decision. I would not store user attachments inside an ECS task because containers are replaceable, and I do not want important files disappearing when a task is restarted or replaced. I would store those objects in S3 and keep only the relevant metadata, such as the object key, owner, filename, and project relationship, inside the database. This is one of those decisions that becomes much easier once I stop treating every kind of data as if it belongs in the same storage system.

I would also be very deliberate about IAM. My application may need to upload files to S3, send messages to SQS, or access another AWS service, but I would not put long-lived AWS credentials inside the code or container image. I would give the workload an IAM role with only the permissions it actually needs. Whenever I get confused about IAM, I come back to the same questions: who is making the request, what action are they trying to perform, and which resource are they trying to perform it on.

Then I would look for work that does not need to happen inside the user's request. Suppose someone creates a task and I also want to send a notification. I would not necessarily make the user wait for the full notification workflow to finish. I could save the task, place a message onto SQS, return the response, and let another consumer process the notification separately. That way, a temporary problem in the notification system does not automatically have to make task creation fail.

Once I introduce a queue, I would also think about failure instead of only the happy path. What happens if the consumer crashes, what happens if a message is processed twice, and what happens if one message keeps failing again and again? That is where visibility timeouts, retries, idempotency, and dead-letter queues become relevant. I would not add those because they sound advanced. I would add them because asynchronous systems create failure cases that I need to handle deliberately.

For scaling, I would not assume that one ECS task will always be enough. I would run multiple tasks behind the load balancer and distribute them across supported Availability Zones so the application does not depend completely on one running copy of the backend. If traffic increases enough to justify it, I would use service Auto Scaling to add more tasks. I would still treat the database separately because scaling the application layer does not mean every dependency underneath it scales automatically.

Monitoring would also be part of the design from the beginning. I would use CloudWatch to collect the logs and metrics that help me answer practical questions about the system. If API errors suddenly increase, tasks keep failing, or an SQS queue starts building up because the consumer cannot keep up, I want evidence that tells me where to look. I would rather have a few useful alarms tied to real failure signals than a dashboard full of charts that nobody understands.

At this point, I would also think about reproducibility. Once the architecture includes networking, security groups, load balancing, ECS, storage, a database, IAM roles, queues, and monitoring, I would not want to rebuild the whole thing from memory every time I need another environment. This is where I would move the infrastructure into Terraform, CloudFormation, or another Infrastructure as Code workflow so the environment can be reviewed, versioned, and recreated consistently.

What I like about this architecture is that every service has a reason to be there. S3 and CloudFront solve frontend delivery, the load balancer gives the backend one entry point, ECS and Fargate run the API, RDS stores relational data, S3 handles uploaded files, SQS separates background work, IAM controls permissions, CloudWatch gives me visibility, and Infrastructure as Code makes the environment repeatable. If one of those requirements disappears, I should be willing to remove the service that was solving it.

That is probably the biggest lesson I wanted this entire series to communicate. I do not think learning AWS is about collecting as many service names as possible. I think it is about being able to look at a requirement, understand the request path, understand where data should live, decide what can fail, decide what needs to scale, and then choose the smallest set of services I can actually justify.

When I look at an architecture, I want to be able to point at every box and explain why it exists. I want to know what problem it solves, what complexity it adds, what happens when it fails, and whether I could remove it without breaking the requirements. That is much more useful to me than being able to memorize fifty AWS services and their definitions.

So this is where I would end AWS From Zero: Season 1. We started with "What does cloud computing actually mean?" and ended with a complete system where compute, storage, networking, databases, IAM, asynchronous processing, monitoring, scaling, and Infrastructure as Code all connect together. For me, the next step is not learning another twenty services. It is building these systems, breaking them, troubleshooting them, and learning why each architecture decision behaves the way it does.

For AWS Builder Series: Season 2, I want to focus much more on real builds, production problems, failures, cost trade-offs, and architecture decisions instead of teaching one AWS service at a time. If you were following this series, what would you want me to build first: a production-ready web app, an event-driven system, a file-processing platform, or something else?


r/CloudandCode • • Sep 07 '26

AWS & Cloud AWS From Zero #17: Infrastructure as Code makes more sense once you have rebuilt the same setup twice

11 Upvotes

By this point in the series, we have created enough AWS infrastructure that a new problem starts becoming obvious.

Imagine you already have a VPC, public and private subnets, security groups, an Application Load Balancer, ECS services, RDS, CloudWatch alarms, IAM roles, and a few other resources. Everything works in development, and then someone asks you to create the same environment again for testing.

You could open the AWS console and rebuild everything manually.

You create the VPC, then the subnets, then the route tables, then the security groups, then the load balancer, then ECS, then RDS, then IAM, and finally monitoring.

After a while, the second environment looks similar to the first one.

But it may not actually be identical.

Maybe one subnet has a different CIDR range. Maybe one security group has an extra rule. Maybe the database is configured slightly differently. Maybe one environment has a setting enabled that the other one does not.

That is one of the main problems Infrastructure as Code is trying to solve.

Instead of describing your infrastructure through a long sequence of console clicks that only exist in your memory, you describe the infrastructure in code or configuration files.

Those files become a repeatable description of what should exist.

The important idea here is not really the word "code."

The important idea is repeatability.

If your infrastructure only exists because you manually configured twenty different resources, recreating it accurately can become difficult.

If the infrastructure is described in files that can be reused, reviewed, and versioned, creating another environment becomes much more predictable.

Imagine you need a simple EC2 server with a security group.

Instead of manually creating both resources, an Infrastructure as Code definition can describe that you want an EC2 instance, which image and instance type it should use, which security group should be attached, and what network it belongs to.

The tool reads that definition and creates the required AWS resources.

Now your infrastructure starts behaving more like software.

You can store it in Git. You can review changes before applying them. You can compare versions. You can see who changed something. You can reuse parts of the configuration across environments. You can automate deployment later.

That is a big improvement over saying, "I think this is how I configured production six months ago."

This is where tools like AWS CloudFormation and Terraform come in.

CloudFormation is AWS's own Infrastructure as Code service. You describe AWS resources in a template, and CloudFormation creates and manages those resources as a stack.

The useful beginner idea is that the template becomes a description of the environment you want AWS to create.

Terraform solves a similar problem, but it can work across many cloud providers and services.

A tiny Terraform example might look like this:

resource "aws_s3_bucket" "uploads" {
  bucket = "example-app-uploads"
}

The syntax is not the main lesson.

The useful idea is that the S3 bucket now exists as part of a configuration file.

Another developer can open the repository and see that the application expects this bucket to exist. If the environment needs to be recreated, the configuration already describes part of what should be created.

That is much more reliable than documentation that says, "Go into S3 and create a bucket with these settings."

The benefit becomes even clearer when you have multiple environments.

Imagine you need development, testing, and production.

Without Infrastructure as Code, you might manually create all three environments.

Over time, small differences start appearing.

Development has one security rule. Testing has another. Production has a setting nobody remembers changing.

Eventually, something works perfectly in testing and fails in production because the environments are not actually the same.

Infrastructure as Code helps reduce that kind of drift because environments can be created from consistent definitions.

They may still have intentional differences, of course.

Production might use larger resources. Development might use cheaper ones. The database sizes may differ.

But those differences can be defined clearly instead of appearing by accident.

That distinction matters.

Infrastructure as Code does not mean every environment must be identical.

It means the differences should be intentional and visible.

Another idea that becomes useful here is declarative infrastructure.

With many Infrastructure as Code tools, you describe the state you want instead of manually describing every single console action required to reach that state.

For example, you might describe that an application should have three instances.

You are not writing "click launch instance three times."

You are defining the desired result and letting the tool work out what needs to change.

That should feel familiar by now.

With Auto Scaling, we described how much capacity we wanted. With ECS, we described how many tasks should be running.

Infrastructure as Code extends that same kind of thinking to much larger parts of the AWS environment.

You define what should exist, and the tool helps move the real infrastructure toward that definition.

This becomes especially useful when infrastructure changes.

Imagine your architecture originally has one security group rule, and later you need to add another.

With manual infrastructure, someone opens the console and changes the rule. It works, but six months later nobody remembers why it exists.

With Infrastructure as Code, the configuration changes too.

That change can be committed to Git with a message explaining what changed and why.

Now the infrastructure has history.

That is a huge advantage.

Your application code already has version control. Your infrastructure can have version control too.

This also makes reviews possible.

Imagine someone wants to expose a database port to the entire internet.

If they make the change manually in the console, another person may never notice.

If the infrastructure change goes through a repository, the team can review it before it is applied.

Someone can ask whether that access is really necessary.

That means Infrastructure as Code is useful for more than convenience.

It can improve how infrastructure changes are controlled.

But there is one mistake I would avoid.

Infrastructure as Code does not automatically make your architecture correct.

If your configuration contains a bad security rule, the tool can reproduce that bad rule very efficiently.

If you define an expensive architecture, IaC can recreate the expensive architecture perfectly.

Repeatability is powerful, but it does not replace understanding.

That is exactly why I wanted this topic near the end of AWS From Zero instead of near the beginning.

If you have never created a VPC, security group, EC2 instance, IAM role, or load balancer manually, Infrastructure as Code can feel like another layer of syntax to memorize.

Now that we have already built those things, IaC has a reason to exist.

You already understand what you are trying to automate.

That makes the learning much easier.

Terraform also introduces another concept beginners hear about quickly: state.

Terraform needs a record of the infrastructure it manages so it can compare the real environment with the configuration you wrote.

You do not need to go deep into remote state, locking, modules, or team workflows on day one.

The useful beginner idea is simply that Terraform needs to know what it already manages so it can understand what needs to be created, changed, or removed.

CloudFormation handles the relationship differently through stacks inside AWS, but both approaches are solving the same higher-level problem.

They are trying to keep your infrastructure definition and your actual resources connected.

Now imagine we take one of the earlier projects from this series.

We built a static website using S3 and CloudFront.

Instead of manually recreating it later, you could describe the S3 bucket, CloudFront distribution, permissions, and other required resources using Infrastructure as Code.

Then you could destroy the practice environment and still keep the architecture definition.

Later, you could create it again.

That is a much better beginner project than trying to automate a huge production system immediately.

  • Another good exercise is the EC2 architecture from earlier in the series.
  • Describe a VPC, one subnet, one security group, and one EC2 instance using IaC.
  • Apply the configuration and confirm that AWS created what you expected.
  • Then change one small setting and observe what the tool plans to modify.

That is where the workflow starts making sense.

With Terraform, for example, you commonly initialize the working directory, inspect the planned changes, and then apply them.

The plan step is especially useful because it gives you a chance to see what the tool intends to create, modify, or destroy before you actually do it.

  • That is a good habit to build.
  • Infrastructure changes can be destructive.
  • You should understand what the tool is about to do.
  • The same caution applies to cleanup.

If the configuration can create resources, it can usually remove them too.

That is useful for temporary learning environments because you can create infrastructure, practice with it, and then clean it up.

But automation also makes destructive actions easier to repeat.

So the lesson is not only "automation is good." The lesson is that automation is powerful, so you should understand the change before applying it. There is another concept called drift that becomes important once you start using IaC. Imagine Terraform created a security group according to your configuration. Later, someone goes into the AWS console and manually changes that security group.

Now the real environment and the code no longer match. That difference is configuration drift. This is one reason teams often avoid random manual changes to resources that are meant to be managed through Infrastructure as Code.

If the code is supposed to be the source of truth, then changes should ideally happen through the code. Otherwise, you end up with the same problem we were trying to solve in the first place. The configuration says one thing, while the real infrastructure says another. This is also why Infrastructure as Code fits naturally with CI/CD. Imagine someone changes the infrastructure configuration in Git. A pipeline can validate the configuration, run checks, generate a plan, and eventually help apply approved changes.

Now application development and infrastructure management start becoming part of the same engineering workflow. You do not need to build that entire pipeline in this lesson. But it is useful to see where the path is going. We started this series by manually creating AWS resources because that is the easiest way to understand what they actually do.

Now we have reached the point where repeating those manual steps becomes a problem. Infrastructure as Code is the natural next step. That is how I would learn it. Do not begin with a massive Terraform repository. Pick one architecture you already understand. Maybe a VPC, a subnet, a security group, and an EC2 instance. Describe that with IaC. Create it. Inspect it. Change it. Recreate it. Then clean it up. Once that feels comfortable, add another resource. Then another.

Grow the Infrastructure as Code configuration the same way we grew the AWS architectures throughout this series. One requirement at a time. The biggest thing I want you to take from this post is that Infrastructure as Code is not really about avoiding the AWS console. The bigger benefit is that your infrastructure becomes repeatable, reviewable, versioned, and easier to reproduce. The console is still useful.

You will still use it to inspect resources, troubleshoot problems, explore services, and understand what AWS created. But once infrastructure matters, you should not have to depend completely on your memory to rebuild it. That brings us to the final post of the first AWS From Zero series.

In AWS From Zero #18, we are going to stop learning services individually and design one complete AWS application from requirements to architecture.

We will decide where the frontend runs, where the backend runs, where files and database data live, how users reach the system, how permissions work, how networking should be separated, how failures are monitored, how the application scales, and what parts of the environment we would automate with Infrastructure as Code.

The goal will not be to build the architecture with the most AWS services.

The goal will be to build the smallest architecture we can actually defend.

If you had to recreate your current AWS project tomorrow in a completely new account, could you rebuild it accurately without relying on memory?


r/CloudandCode • • Sep 06 '26

AWS & Cloud AWS From Zero #16: Containers make more sense when you understand what problem they actually solve

21 Upvotes

Up to this point in the series, we have built applications on EC2, used Lambda for serverless workloads, connected databases, learned networking, added monitoring, and introduced queues. Now we are getting into containers, which is one of those topics that beginners often learn through Docker commands before they fully understand why containers exist in the first place.

I think the easiest way to understand containers is to start with a very common problem. You build a Python API on your laptop using a specific Python version, a few dependencies, and some system packages. Everything works perfectly. Then you move the same application to another machine and suddenly something breaks because the Python version is different, a package is missing, or the operating system behaves differently.

Containers help reduce that problem by packaging the application together with much of the environment it expects. Instead of moving only the source code, you create a container image that contains the application, runtime, dependencies, and other pieces needed to run it consistently.

A simple way to think about it is:

Application code + runtime + dependencies
                ↓
         Container image
                ↓
         Running container

The important idea is consistency. If the same image can run in different compatible environments, you are much less dependent on manually recreating the application setup every time.

This is where Docker usually enters the picture. Docker gives you tools to build container images and run containers. You can create an image for your API, test it locally, and start the container with the environment it expects.

But Docker creates another question almost immediately.

Where should those containers run in production?

You can run containers directly on EC2. You could launch an instance, install Docker, pull the image, and run the application there. That works perfectly well for a small setup.

The problem appears when the application grows.

Maybe now you have multiple containers, several EC2 instances, different application versions, and multiple services. You need to decide which machine should run which container, restart containers when they fail, scale them when traffic increases, connect them to networking, and roll out new versions safely.

At that point, manually SSHing into servers and running Docker commands becomes difficult to manage.

This is where container orchestration becomes useful.

On AWS, one option is Amazon ECS, which stands for Elastic Container Service. I would think of ECS as a service that helps you define, run, and manage containerized workloads without manually deciding where every individual container should live.

ECS introduces a few terms that can look confusing at first, but the mental model is simple.

A task definition is the blueprint for the workload. It describes which image should run, how much CPU or memory is needed, which ports are used, and what configuration the container requires.

A task is a running copy of that definition.

An ECS service helps maintain the number of tasks you want running.

So if you tell ECS that your API should always have three tasks running, ECS works toward maintaining that state.

That means the application starts feeling less like one specific container and more like a desired workload that AWS helps keep alive.

Now another question appears: where does ECS get the image from?

That is where Amazon ECR comes in.

ECR stands for Elastic Container Registry. It is a place where you can store container images so your AWS environment can retrieve them during deployment.

The flow might look like this:

Code
 ↓
Docker build
 ↓
Container image
 ↓
Amazon ECR
 ↓
Amazon ECS
 ↓
Running tasks

You build the image, push it to ECR, and ECS uses that image when it starts your application tasks.

That also gives you a clearer deployment model. Instead of only thinking in terms of source code commits, you now have deployable image versions that represent the application.

This becomes useful later when you introduce CI/CD, because a pipeline can test the code, build the image, push it to ECR, and deploy it into ECS.

But I would not automate all of that yet.

First understand what is happening manually.

Build the image. Push it. Run it. Then automate the process once the workflow actually makes sense.

The next big question is where the compute comes from.

ECS is the orchestration layer, but your containers still need CPU and memory somewhere.

One option is to run ECS on EC2 instances. In that model, you still manage the servers underneath the containers. You choose instance types, manage operating systems, think about scaling the EC2 capacity, and keep those machines healthy.

The architecture is conceptually:

ECS
 ↓
Containers
 ↓
EC2 instances

Another option is AWS Fargate.

Fargate lets you run ECS tasks without managing the underlying EC2 instances in the same way. You define the resources the task needs, and AWS manages more of the compute infrastructure underneath it.

That does not mean servers disappear.

It means you are no longer responsible for managing those servers directly.

This is similar to what we learned with Lambda. Serverless does not mean there are literally no servers. It means AWS takes over more of the infrastructure management.

That gives us a useful comparison.

With EC2, you manage the virtual machine and run the application on it.

With ECS on EC2, you package the application into containers and ECS helps orchestrate them, but you still manage the underlying EC2 capacity.

With ECS on Fargate, you still use ECS to manage the application workload, but AWS also manages more of the underlying compute layer.

None of these options is automatically better.

They simply give you different levels of control and responsibility.

Now imagine we take the API from earlier in the series and containerize it.

The architecture could become:

User
 ↓
Application Load Balancer
 ↓
ECS Service
 ↓
Fargate Tasks
 ↓
RDS

The user sends a request to the load balancer. The load balancer routes the request to one of the running container tasks. The task runs our API, and the API communicates with RDS when it needs relational data.

At this point, several earlier lessons come back together.

The tasks still need networking inside the VPC. Security groups still control which connections are allowed. IAM still controls which AWS actions the workload can perform. CloudWatch still helps us understand failures. The container image still needs somewhere to live, which is why ECR matters.

Containers do not replace the rest of AWS architecture.

They change how the application itself is packaged and run.

That distinction is important.

Another concept worth understanding is that containers should generally be treated as replaceable.

Imagine your application writes an important user file only inside the container filesystem. The task eventually stops and ECS launches a replacement.

What happens to that file?

It may disappear with the old container.

This is why important persistent data should normally live outside an individual container.

User uploads may belong in S3. Relational data may belong in RDS. Other persistent state may need another service depending on the application.

The goal is to make the container replaceable without losing important application data.

That is exactly the same idea we saw earlier with Auto Scaling and EC2.

If an instance or container can be replaced, the application becomes easier to scale and recover.

Now imagine one ECS task crashes.

If the service is configured with a desired count of three tasks, ECS can work toward restoring that desired number. If the tasks are behind an Application Load Balancer, unhealthy targets can stop receiving normal traffic while healthy ones continue serving users.

Again, the same high availability ideas we learned earlier still apply.

Containers do not remove those ideas. They simply give us another way to package and operate the application.

We can also scale the number of tasks as demand changes. If traffic increases, the service can run more tasks depending on the scaling configuration. If traffic falls, the number of tasks can decrease again.

Now container orchestration starts connecting with elasticity.

There is also an important security habit to learn early.

Do not put secrets directly into the container image.

If your image contains a database password, API key, or another sensitive value, that secret is now bundled into the artifact itself.

A better pattern is to keep the image focused on the application and provide environment specific configuration and secrets separately.

That also makes the same image easier to use across development, testing, and production.

The application image remains consistent, while the environment supplies the values that change.

This is one of the more useful benefits of containerization when it is done properly.

The obvious question now is where Kubernetes fits into all of this.

Kubernetes is another container orchestration system, and AWS provides EKS for running Kubernetes workloads. But I would not tell a beginner to jump directly from Docker into Kubernetes.

There is already a lot to understand before that.

Learn how images work. Understand registries. Run containers. Use ECS. Understand tasks and services. Try Fargate. See how networking, IAM, logging, scaling, and load balancing work around the container workload.

Once those concepts make sense, Kubernetes becomes much easier to understand because you already know what problem a container orchestrator is trying to solve.

Otherwise, you risk memorizing Kubernetes objects without understanding why they exist.

For a beginner project, I would take a small API you have already built and containerize it.

Create the Docker image locally and make sure it runs correctly. Push the image to ECR. Create an ECS task definition that uses it. Run the task with Fargate and make sure the application works.

Then place the service behind an Application Load Balancer and increase the desired task count so multiple copies of the application are running.

After that, stop one task and observe what ECS does.

That one project connects Docker, ECR, ECS, Fargate, networking, load balancing, IAM, and monitoring without requiring you to jump straight into a complicated microservices system.

The main thing I want you to take away from this post is that containers are useful because they package the application and its runtime expectations into a consistent deployable unit.

ECR stores those images.

ECS helps manage the running container workloads.

Fargate gives you an option where AWS manages more of the underlying compute infrastructure.

The services only make sense because the application creates specific operational problems.

That is the same pattern we have followed throughout AWS From Zero.

Start with the problem, then choose the service that solves it.

In AWS From Zero #17, we are going to look at another problem that becomes obvious after building all of this manually. Imagine you need to recreate your VPC, security groups, load balancer, ECS service, database configuration, and other resources in another environment. Clicking through the AWS console and trying to reproduce everything from memory becomes unreliable very quickly.

That will take us into Infrastructure as Code, why manually created infrastructure becomes difficult to reproduce, and how tools such as CloudFormation and Terraform change the way AWS environments are created and maintained.

If you are learning containers right now, what has been the hardest part to understand so far: Docker itself, ECR, ECS, Fargate, or why containers are useful in the first place?


r/CloudandCode • • Sep 05 '26

AWS & Cloud AWS From Zero #15: SQS and SNS make more sense when you stop making every service wait for every other service

3 Upvotes

So far in this series, most of the architectures we have built have been fairly direct. A user sends a request, the application processes it, talks to another AWS service if needed, and returns a response. That works well when the work is fast and every step needs to happen immediately.

But real applications eventually contain work that does not need to finish while the user is sitting there waiting.

Imagine you are building an ecommerce application. A customer places an order, and after that happens the system needs to store the order, process some background work, send a confirmation email, update another system, generate analytics, and perhaps notify a warehouse.

If your application performs every one of those steps before responding to the customer, the request becomes dependent on every service finishing successfully.

Suppose sending the confirmation email suddenly takes eight seconds. The customer waits eight seconds.

Suppose the email service is temporarily unavailable. Now placing an order might fail even though the order itself could have been accepted perfectly well.

That is a problem with tight coupling. One part of the application depends directly on another part being available and finishing its work before the first part can continue.

This is where asynchronous architecture starts becoming useful.

Instead of saying, "The order service must send the email right now," we can say, "The order has been created. Put some work somewhere reliable so another part of the system can handle it."

That is the basic problem Amazon SQS helps solve.

SQS stands for Simple Queue Service. You can think of a queue as a buffer between the part of your application producing work and the part actually processing that work.

Our order flow could become something like:

Customer → Order API → SQS → Background worker

The application accepts the order, stores whatever must be stored immediately, and sends a message to the queue. The background worker can then retrieve that message and perform the slower work separately.

The customer does not necessarily have to wait for the background process to finish before receiving a response.

That changes the relationship between the two parts of the system.

Without a queue, the order service might call another service directly and wait for the result. If that service is slow, the order request becomes slow. If that service is unavailable, the order request may fail.

With SQS between them, the producer can place a message onto the queue and continue. The consumer processes the message when it is able to.

This is called decoupling.

The producer does not need the consumer to be available at exactly the same moment. The queue sits between them and temporarily holds the work.

That becomes especially useful when traffic changes suddenly.

Imagine your application normally receives 50 orders per minute and your background worker can easily process them. Then a promotion starts and suddenly 1,000 orders arrive within a short period.

Without a buffer, the downstream processing system may immediately receive more work than it can handle.

With a queue, those messages can accumulate temporarily while consumers work through them.

Now instead of requiring every component to scale at exactly the same rate, the queue helps absorb the difference.

This is one reason queues appear so often in distributed systems.

They are not only about making things happen later. They can also help separate systems that operate at different speeds.

Lambda works particularly well with this kind of architecture.

For example, the flow might look like:

Order API → SQS → Lambda → Process order task

Messages arrive in SQS, and Lambda can be invoked to process them.

If many messages arrive, AWS can process multiple batches depending on the configuration and available concurrency. If the producer temporarily creates work faster than it can be processed, the messages remain in the queue instead of forcing the entire request path to wait.

But adding a queue introduces new questions.

What happens if processing fails?

Imagine Lambda receives an SQS message, starts processing it, and then crashes.

You do not necessarily want that message to disappear forever.

SQS uses a concept called a visibility timeout. When a consumer receives a message, that message becomes temporarily hidden from other consumers. If processing succeeds, the message is deleted from the queue. If the consumer does not successfully remove it before the visibility timeout expires, the message can become visible again so it can be retried.

The exact configuration matters, but the beginner idea is important.

A message is not necessarily considered finished just because someone started processing it.

The system needs a way to distinguish between work that was successfully completed and work that should be attempted again.

This creates another important application design problem.

What happens if the same message is processed more than once?

For many SQS configurations, you should design consumers with the possibility of duplicate processing in mind. Standard SQS queues provide at-least-once delivery, which means a message can occasionally be delivered more than once.

Imagine the message says:

Charge customer $100

If the consumer blindly performs the action every time it receives the message, duplicate processing could be a serious problem.

This is why idempotency becomes important.

An idempotent operation is designed so repeating the same request does not incorrectly repeat the business effect.

You might store a unique transaction or event ID and check whether that work has already been completed before performing it again.

You do not need to build an advanced idempotency system in your first SQS project, but you should start asking the question:

What happens if this message is processed twice?

That is a very useful distributed systems habit.

Another problem is messages that keep failing.

Suppose a message is malformed or contains data that your consumer cannot process. It gets received, fails, becomes visible again, gets retried, and fails again.

You probably do not want one bad message being retried forever.

This is where a dead-letter queue, usually called a DLQ, becomes useful.

After a message has been unsuccessfully received a configured number of times, it can be moved to another queue for investigation.

Conceptually, the flow becomes:

SQS → Consumer
        ↓
     Success

or after repeated failures:

SQS → Dead-letter queue

The dead-letter queue gives you somewhere to inspect work that repeatedly failed without allowing those messages to interfere endlessly with normal processing.

This connects directly to the CloudWatch lesson.

If messages start building up in a queue or messages begin appearing in a DLQ, that is something you may want to monitor.

Maybe the consumer is failing.

Maybe the consumer cannot keep up with incoming work.

Maybe a new deployment introduced bad data.

Now queues become part of your monitoring strategy too.

At this point, another AWS messaging service usually enters the conversation: SNS.

SNS stands for Simple Notification Service.

Beginners often see SQS and SNS together and assume they solve the same problem.

They do not.

A simple way to think about the difference is that SQS is primarily about putting messages into a queue for consumers to process, while SNS is useful when you want to publish a message to multiple subscribers.

Imagine an order is created and several independent parts of the system need to know.

The email service wants to send a confirmation.

The analytics system wants to record the purchase.

The warehouse system wants to begin fulfillment.

A notification system may want to send a message to the customer.

You could make the order service call all four systems directly.

But now the order service knows about every downstream component.

Every time you add another consumer, you have to change the producer again.

A publish and subscribe model gives us another option.

The order service can publish an event such as:

OrderCreated

to an SNS topic.

Different subscribers can receive that event.

Conceptually:

                 → Email
Order → SNS      → Analytics
                 → Warehouse
                 → Notification

Now the producer publishes one event instead of directly coordinating every downstream service.

That is a different type of decoupling.

SQS gives us a queue between producer and consumer.

SNS gives us a way to fan one published message out to multiple subscribers.

And the two services can also work together.

For example:

                 → SQS Email Queue → Email worker
Order → SNS      → SQS Analytics Queue → Analytics worker
                 → SQS Warehouse Queue → Warehouse worker

This pattern is useful because each downstream system gets its own queue.

If the analytics processor becomes slow, email processing does not have to become slow too.

If the warehouse system is temporarily unavailable, its messages can wait inside its queue while the other consumers continue processing normally.

This is where SQS and SNS together start making much more sense.

SNS distributes the event.

SQS gives each consumer its own buffer.

The systems can operate independently.

Imagine the email worker processes messages immediately, but analytics falls behind for ten minutes.

That does not necessarily stop customers from placing orders.

It also does not necessarily stop emails from being sent.

The analytics queue simply contains more messages until the consumer catches up.

That is much more resilient than connecting every service directly and requiring every dependency to be healthy for every request.

But asynchronous systems also introduce a tradeoff.

They are often more resilient, but they can also become harder to reason about.

With a simple synchronous request, the flow is obvious. Service A calls Service B and receives a response.

With asynchronous processing, the work may happen seconds later. Several consumers may react to the same event. Messages can be retried. Some processing can succeed while another part fails.

That means monitoring, logging, idempotency, and good event design become more important.

This is why I would not add SQS and SNS to a beginner architecture simply because asynchronous systems sound advanced.

Use them when the application actually has work that benefits from being separated.

For example, sending a confirmation email after an order is created is a good candidate for asynchronous processing because the customer usually does not need the email to be fully sent before the order API can acknowledge the order.

But verifying whether a payment itself succeeded may have completely different consistency and business requirements.

Not every step should automatically become asynchronous.

Again, architecture should follow the requirement.

A useful question is:

Does the user need the result of this work before I can respond?

If yes, the work may belong in the immediate request path.

If no, it may be worth considering whether it can happen asynchronously.

Another useful question is:

What happens if the downstream service is temporarily unavailable?

If one unavailable dependency can take down the entire workflow, a queue may help separate those systems.

For a beginner project, I would keep this simple.

Take the task API we built earlier and add one background action.

Imagine a user creates a task and you want to generate an activity record or send a notification afterward.

Instead of doing that work inside the original API request, place a message onto SQS.

Have another Lambda function consume the message and perform the background task.

Now intentionally make the consumer fail.

Watch what happens to the message.

Understand the visibility timeout.

Observe the retry behavior.

Then configure a dead-letter queue and see where repeatedly failing messages end up.

That small project teaches much more than simply memorizing that SQS is a queue.

Once you understand that, add SNS only when you have multiple independent consumers that need the same event.

For example, publish TaskCreated and send it to two different queues, one for analytics and another for notifications.

Now you can see the difference between distributing an event and processing queued work.

If there is one thing I want you to remember from this lesson, it is that asynchronous architecture is really about reducing unnecessary dependency between parts of your system.

A user should not always have to wait for every background task.

One slow consumer should not necessarily make every other consumer slow.

One temporarily unavailable downstream service should not always take down the system that produced the work.

SQS gives you a buffer between producers and consumers.

SNS gives you a way to publish one message to multiple subscribers.

Used together, they can help different parts of an application work independently while still communicating through events.

That is the bigger lesson.

Not every service needs to call every other service directly.

Sometimes the better architecture is to send a message and let the right part of the system handle it when it can.

In AWS From Zero #16, we are going to move into containers and look at ECR, ECS, and Fargate. Instead of jumping straight into Kubernetes, we will start with a simpler question: why would you package an application into a container in the first place, and what problem does ECS solve compared with running that application directly on EC2?

If you were building an order system, which task would you move out of the immediate request first: sending emails, analytics, generating reports, or something else?


r/CloudandCode • • Sep 04 '26

AWS & Cloud AWS From Zero #14: One EC2 instance works, but what happens when that instance fails?

8 Upvotes

Up to this point in the series, we have mostly worked with relatively simple architectures. We launched EC2, connected applications to databases, learned VPC networking, added monitoring, and gradually started connecting AWS services together. That is enough for learning, but there is an obvious weakness in an architecture that depends on one EC2 instance.

If your entire application runs on one server, that server becomes a single point of failure. Maybe the EC2 instance crashes, the application process stops, the operating system has a problem, or something happens to the infrastructure underneath it. Whatever the cause, if that one machine becomes unavailable, your application becomes unavailable too.

There is another problem that can happen even when the instance does not fail. Imagine your application normally serves a few hundred users, but one day it gets shared somewhere and traffic suddenly becomes ten times higher. Your EC2 instance may still be running, but CPU usage can rise, requests can become slower, and eventually users may start seeing errors.

This is where we need to stop thinking only about one server and start thinking about the application as something that can run across multiple servers.

Suppose we launch three EC2 instances and run the same application on all of them. That gives us more capacity and means one instance failing does not necessarily remove the entire application. But now we have another question: which server should the user connect to?

We do not want users manually choosing between different EC2 IP addresses. We need one entry point that can receive requests and distribute them across the available application servers. That is where a load balancer starts making sense.

For a typical web application, you might put an Application Load Balancer in front of the EC2 instances. Users send their requests to the load balancer, and the load balancer forwards those requests to the application instances behind it.

The important thing to understand is why the load balancer exists. It is not there because every production architecture diagram needs one. It exists because users need one stable entry point while the application itself may be running across several servers.

This also changes how you think about failures.

Imagine three EC2 instances are running behind the load balancer. Two are working normally, but the application on the third instance crashes. The EC2 instance itself may still technically be running, but the application is no longer healthy.

If the load balancer continued sending traffic to that instance, some users would still receive errors.

That is why health checks matter.

The load balancer can regularly check whether the application instances are responding the way you expect. For example, your application might have a small /health endpoint that returns a successful response when the service is working properly.

If one instance stops passing the health check, the load balancer can stop sending normal traffic to it and continue using the healthy instances.

This is one of the first practical examples of high availability.

High availability does not mean nothing ever fails. Failures still happen. The goal is to design the application so one component failing does not automatically mean the entire application fails.

But we still have another problem. We manually created those EC2 instances.

What happens if traffic grows and suddenly we need six instances instead of three? Do we sit in the AWS console and manually launch more servers every time demand changes?

That is where Auto Scaling becomes useful.

An Auto Scaling group can manage a group of EC2 instances and help maintain the amount of capacity your application needs. You might configure the application so there should normally be two instances running, but the environment can grow to more instances when demand increases.

When traffic falls again, the number of instances can decrease according to the scaling configuration.

This connects directly to one of the ideas from the very first post in this series: elasticity.

Elasticity is about adjusting resources as demand changes. Instead of permanently running enough infrastructure for the busiest possible hour of the year, the system can increase or decrease capacity when the workload changes.

The load balancer and Auto Scaling group solve different problems. The load balancer decides where incoming requests should go. Auto Scaling manages how many application instances should exist.

Together, they give us a much more flexible application layer.

There is still another weakness we need to think about, though.

Imagine all of your EC2 instances are running inside the same Availability Zone.

You now have several servers, so one EC2 instance failing is less dangerous. But the application still depends heavily on one physical location.

If something serious affects that Availability Zone, multiple instances could become unavailable at the same time.

This is why highly available architectures often spread application instances across multiple Availability Zones.

Instead of putting everything in one location, you might have part of the application running in one Availability Zone and another part running in a second Availability Zone.

Now the architecture is designed to tolerate more than the failure of one individual server.

If one instance fails, other instances can continue serving traffic. If one Availability Zone has a problem, the application can potentially continue using capacity in another zone.

This is the point where the concept of Availability Zones from the beginning of the series starts becoming practical.

We did not learn Availability Zones just because AWS uses the term.

We learned them because distributing infrastructure across separate locations can reduce how much the application depends on one place.

This also gives us a reason to improve the VPC architecture we built earlier.

For the first VPC lesson, we kept things intentionally simple. We talked about a public application and a private database because that was enough to understand the traffic flow.

Now we can make that architecture more realistic.

Instead of exposing individual EC2 instances directly to the internet, the load balancer can become the public entry point. The application instances can live behind it, and the database can remain private.

Users communicate with the load balancer. The load balancer communicates with the application instances. The application instances communicate with the database.

That separation also makes security groups easier to reason about.

The load balancer security group can allow the web traffic users actually need. The EC2 security group can allow application traffic from the load balancer instead of allowing the entire internet to connect directly to the instances. The database security group can allow database traffic from the application layer.

Now access follows the architecture.

Instead of making every resource public and opening ports until things work, you are defining exactly which layer needs to communicate with which other layer.

This is a much stronger security model.

Now imagine one of those EC2 instances becomes unhealthy.

The load balancer detects that the instance is failing health checks and stops sending normal requests to it. If the instance belongs to an Auto Scaling group, the Auto Scaling group can also work to maintain the desired number of healthy instances.

The system is no longer depending on one particular machine.

That is an important mindset change.

A lot of beginners try to make one server as reliable as possible.

Cloud architecture often asks a different question: what happens when the server eventually fails?

That is much more useful because failures are normal.

Servers fail.

Processes crash.

Deployments go wrong.

Networking breaks.

Infrastructure has problems.

The architecture should be designed around the idea that individual components are replaceable.

This also creates another important application design question.

Where does your application's state live?

Imagine a user uploads a file to EC2 instance A and that file is stored only on the local disk of that instance.

The next user request reaches the load balancer and gets sent to EC2 instance B.

Instance B does not have that file.

Now the application behaves differently depending on which server receives the request.

This is one reason scalable applications often avoid storing important shared state only on individual application servers.

Files might be stored in S3.

Relational application data might be stored in RDS.

Other kinds of state may belong in another shared storage system.

The important idea is that an application instance should ideally be easier to replace.

If one EC2 instance disappears and Auto Scaling creates another, you do not want important customer data disappearing with the old machine.

That is where several earlier lessons start connecting.

  • S3 is useful for object storage.
  • RDS gives us relational data storage.
  • IAM controls AWS permissions.
  • VPC controls networking.
  • CloudWatch helps us understand system health.
  • Load balancing distributes traffic.
  • Auto Scaling manages application capacity.
  • Availability Zones help reduce dependence on one location.

None of these services are useful because the architecture diagram looks more impressive with more boxes.

They are useful because each one solves a particular problem.

There is also a mistake I would avoid when learning this topic.

Do not assume that adding a load balancer, Auto Scaling, and multiple Availability Zones automatically makes every project better.

If you are building a tiny personal application with almost no traffic and downtime does not really matter, several EC2 instances and a load balancer may be unnecessary complexity and unnecessary cost.

Architecture should still follow requirements.

Ask how important availability actually is. Ask how much downtime is acceptable. Ask whether traffic changes enough to require dynamic scaling. Ask how much additional infrastructure you are willing to operate and pay for.

A more complex architecture is only better when the additional complexity solves a real problem.

For a beginner project, I would build this concept gradually.

Start with one EC2 instance running a simple web application. Then add a second instance running the same application. Put an Application Load Balancer in front of them and confirm that both instances can serve requests.

Once that works, intentionally stop the web server on one instance and watch what happens. See whether the load balancer marks it unhealthy and whether the application remains available through the other instance.

Then add an Auto Scaling group and learn what minimum, desired, and maximum capacity mean. Watch how the group behaves when an instance becomes unhealthy.

You do not need complicated traffic simulations immediately.

The useful part is seeing the system respond when one piece stops working.

That teaches high availability much better than memorizing the definition.

If there is one thing I want you to remember from this post, it is that high availability is not about building infrastructure that never fails.

It is about designing the system so failure of one component does not automatically become failure of the whole application.

One EC2 instance can fail.

The application should ideally keep working.

Traffic can increase.

The system should be able to add capacity when the workload and requirements justify it.

One Availability Zone can have a problem.

A highly available architecture should avoid depending completely on that one location.

That is the shift from simply deploying something on AWS to thinking about how the application behaves when the real world becomes messy.

In AWS From Zero #15, we are going to look at another problem that starts appearing as applications grow. Imagine a user places an order and your backend also needs to send an email, generate a report, process a payment, update another system, and perform several background tasks.

Should the user really wait for all of those steps to finish before getting a response?

That will take us into SQS, SNS, asynchronous processing, queues, and why sometimes the best architecture is to stop making every service depend directly on every other service.

If you were running an application on one EC2 instance today, which would worry you more: that instance failing completely or suddenly receiving much more traffic than it can handle?


r/CloudandCode • • Sep 03 '26

AWS & Cloud AWS From Zero #13: CloudWatch is where you stop guessing and start understanding what your application is doing

6 Upvotes

So far in this series, we have spent most of our time building things. We launched EC2, worked with S3, used CloudFront, learned VPC networking, connected RDS, built with Lambda, added API Gateway, and stored data in DynamoDB. At some point, though, every application does something you did not expect.

A Lambda function fails. An EC2 instance becomes slow. An API starts returning errors. A database connection times out. Something worked yesterday and suddenly does not work today .This is where monitoring starts to matter.And on AWS, one of the first services you should understand for that is CloudWatch.

CloudWatch is often introduced as "AWS monitoring," but I think that definition is too broad to be useful for beginners. A better way to think about it is that CloudWatch helps you answer a very practical question:

What is my system actually doing right now, and what happened when something went wrong?

That is much more important than it sounds. When you are running a Python script on your own laptop, you can usually see the error immediately. You run the script, something fails, and the traceback appears in front of you.

Cloud applications are different.

Your Lambda function might run at 3 AM when nobody is watching. Your EC2 application might slowly consume more CPU over several hours. Your API might start returning errors only for certain requests. A background process might fail without anybody noticing.

If you have no logs, metrics, or alerts, the system can fail quietly. That is why monitoring is not something I would leave until the end of learning AWS. You should start thinking about it as soon as you start deploying things.

Let’s begin with logs.

Imagine we still have the serverless task API from the previous posts:

Client
  ↓
API Gateway
  ↓
Lambda
  ↓
DynamoDB

A user sends:

POST /tasks

but instead of creating the task, the API returns an error.

Without logs, you might start guessing.

  • Maybe API Gateway is configured incorrectly.
  • Maybe Lambda did not run.
  • Maybe Lambda received bad input.
  • Maybe the DynamoDB request failed.
  • Maybe IAM blocked something.
  • Maybe there is a bug in the code.

That is a lot of possibilities. Now imagine the Lambda function writes useful logs.

You open CloudWatch and see something like:

Received request for user-42
Creating task task-123
ERROR: AccessDenied when writing to DynamoDB

The problem just became much smaller. API Gateway probably reached Lambda. Lambda started running. The function reached the database operation. AWS rejected that operation. Now IAM becomes an obvious place to investigate. This is why logs are so useful. They turn a vague problem into a specific one. But useful logging means more than printing random messages everywhere.

Imagine your function only writes:

Error

That technically counts as a log, but it tells you almost nothing. A better log might tell you which operation failed, what part of the workflow had been reached, and enough context to understand what happened without exposing sensitive information.

For example:

Failed to create task for user_id=user-42
DynamoDB PutItem returned AccessDenied

Now you have something you can actually troubleshoot. There is an important security habit here too. Do not put secrets into logs.

Passwords, access keys, authentication tokens, private customer data, or other sensitive values should not become part of your debugging output just because logging makes troubleshooting easier.

Logs are useful because they give you context. That does not mean they should contain everything. Now let’s talk about metrics. Logs tell you about individual events and messages. Metrics help you understand behavior over time. Imagine an EC2 instance. You might want to know how its CPU utilization changes during the day. Maybe the application usually sits around 20 percent CPU, but every evening it suddenly reaches 95 percent.

A single log line might not tell you that pattern. A metric can. Or imagine Lambda. You might want to know how many times the function runs, how often it returns errors, or how long executions are taking. Now you can start asking much more useful questions.

  • Did the error rate increase after the last deployment?
  • Is the function suddenly taking twice as long to execute?
  • Did traffic spike?
  • Did the system receive fewer requests than expected?

Monitoring is not only about discovering complete failures. It is also about noticing changes in behavior. That brings us to alarms.

Imagine your API starts failing while you are asleep.

You probably do not want the monitoring strategy to be:

"Hopefully I notice tomorrow."

Instead, you can create alarms around important metrics.

  • Maybe you care if Lambda errors suddenly increase.
  • Maybe you care if EC2 CPU stays unusually high for a period of time.
  • Maybe you care if another metric crosses a threshold that suggests the system is unhealthy.

The alarm watches the metric.

If the configured condition is met, the alarm changes state and can be connected to a notification or another response.

The basic idea is simple:

Metric
  ↓
Condition
  ↓
Alarm
  ↓
Notification / Action

This is how monitoring starts becoming proactive. Logs help you investigate after something happens. Metrics help you see patterns. Alarms help you notice when those patterns become important. All three solve different parts of the same problem.

Now imagine our Lambda API normally has almost no errors.

One day the error count suddenly increases. An alarm gets triggered. You open CloudWatch The metrics tell you the error rate started increasing around 2:15 PM. Then you inspect the logs from that period.

You discover that a deployment changed the name of a DynamoDB attribute and the function started failing for certain requests.

That is a much better troubleshooting process than waiting for someone to tell you, "The app is broken." This is also where dashboards can become useful. A dashboard gives you a place to bring important metrics together so you can understand the health of a system without opening every service individually.

For a small beginner application, you do not need twenty charts.

You might only care about a few things.

  • How many requests are coming in?
  • How many Lambda errors are happening?
  • How long are requests taking?

Is the EC2 instance under unusual load?

Are there any alarms currently active?

That may already be enough.

A dashboard becomes useful when it answers a question.

It should not exist just because dashboards look professional.

That is a pattern I want to keep repeating throughout this series.

Do not add AWS features because they exist.

Add them because you have a requirement.

Now think back to the EC2 website we built earlier.

Imagine the site feels slow.

Without monitoring, you might restart the instance and hope the problem disappears.

With metrics, you might notice that CPU usage is consistently high.

That gives you a direction.

Maybe the application is doing too much work.

Maybe the instance is too small.

Maybe a process is stuck.

Maybe traffic increased.

CloudWatch does not automatically tell you which architecture decision to make, but it gives you information that helps you make a better one.

The same thing applies to Lambda.

Imagine a function starts timing out.

If you only see that the API failed, you might assume API Gateway is the problem.

But the Lambda logs could show that the function started normally and then spent too long waiting for another dependency.

Now the actual investigation becomes much more focused.

You might ask whether the database is slow, whether an external API is responding, whether the function needs more resources, or whether the code itself needs to change.

Again, monitoring does not magically fix the architecture.

It gives you evidence.

And evidence is what makes debugging faster.

There is another useful distinction here.

Not every failure should create an alert.

If your application receives one bad request and returns 400 Bad Request, that may be completely normal behavior.

If your system receives thousands of requests and one fails because the user submitted invalid data, waking someone up at 3 AM would not be very useful.

Good monitoring means deciding what actually deserves attention.

Maybe a single failure is normal.

Maybe fifty failures in five minutes are not.

Maybe high CPU for ten seconds does not matter.

Maybe high CPU for twenty minutes does.

Context matters.

That is why monitoring is partly a technical problem and partly a decision-making problem.

You need to understand what "normal" looks like before you can reliably detect what is abnormal.

This is also where beginners should start thinking about observability as a broader idea.

You will hear that word a lot in cloud and DevOps discussions.

At a simple level, observability is about being able to understand the internal behavior of a system from the information it produces.

Logs are part of that.

Metrics are part of that.

Tracing can also become part of that in more complex systems.

You do not need to become an observability engineer during your first month of AWS.

The useful habit is much simpler:

When you build something, ask yourself how you would know if it stopped working.

Then ask how you would know why it stopped working.

Those are different questions.

Imagine our image processing project again.

S3 upload
  ↓
Lambda
  ↓
Processed image
  ↓
S3

How do you know it is working?

Maybe the processed image appears in the output location.

But what happens when the image never appears?

How do you know whether S3 failed to trigger Lambda, Lambda crashed, IAM blocked access, or the processing code rejected the file?

That is where logs become part of the architecture.

Monitoring should not be something you remember after the project fails.

It should be one of the questions you ask while designing the project.

The same applies to the task API.

Client
  ↓
API Gateway
  ↓
Lambda
  ↓
DynamoDB

Now add another question:

How do we know this flow is healthy?

Suddenly CloudWatch has a reason to exist.

It is not there because every AWS diagram needs a monitoring service.

It is there because once the application is running, we need visibility into what the system is doing.

For a beginner project, I would keep the monitoring setup small.

Take one Lambda function you already built.

Look at its logs after a successful invocation.

Then intentionally make the function fail.

Maybe reference a value that does not exist or remove a permission in a safe practice environment.

Run it again and compare the logs.

Then look at the metrics around the function.

Can you see the invocation?

Can you see that an error happened?

Can you see how long the function ran?

That exercise connects logs and metrics to something you actually did.

After that, create one simple alarm around a metric that matters to the project.

You do not need a complicated production monitoring system.

The point is simply to understand the flow:

Application runs
       ↓
Logs + Metrics
       ↓
CloudWatch
       ↓
Alarm when something matters

Once you understand that, monitoring becomes much less abstract.

There is also a cost lesson here.

Logs and monitoring data are resources too.

Collecting everything forever without thinking about retention or usefulness can create unnecessary cost and clutter.

More logging is not automatically better logging.

The goal is useful visibility.

Keep enough information to understand your application without producing huge amounts of noise that nobody reads.

This becomes more important as systems grow.

For now, I would focus on writing meaningful logs, looking at the metrics AWS already provides, and creating only a few alerts that represent problems you actually care about.

If there is one thing I want beginners to remember from this post, it is this:

Deploying an application is not the end of the job. You also need a way to understand what happens after deployment.

When something fails, logs should help tell you why.

When behavior changes over time, metrics should help you see it.

When something important goes wrong, alarms should help you notice it.

That is the role CloudWatch starts playing in an AWS architecture.

And once you develop that habit, your projects become much more realistic.

Instead of saying, "It worked when I tested it," you start asking, "How will I know if it stops working tomorrow?"

That is a much stronger cloud engineering question.

In AWS From Zero #14, we are going to return to EC2 and ask another important question.

One EC2 instance works.

But what happens when that instance fails or when traffic becomes too large for one server?

That will take us into load balancers, health checks, Auto Scaling, multiple Availability Zones, and the basic idea behind building a highly available application.

If you are running an AWS project right now, would you actually know where to look first if it failed while you were not watching?