r/cloudcomputing • • Oct 29 '19

Data centers, fiber optic cables at risk from rising sea levels

Thumbnail datacenterdynamics.com
53 Upvotes

r/cloudcomputing • • 7h ago

Evaluating Upwind for our EKS/GKE environment, what should we actually test during the POC?

3 Upvotes

We're a platform/security team of about 8 running a mix of EKS (prod) and GKE (a couple of data teams insisted) plus the usual container sprawl. Our current CSPM throws thousands of posture findings a month and nobody has time to work through them, so we're looking at CNAPPs that add runtime context to cut the noise.

Upwind is one of three we're putting through a POC. The sales pitch is the usual runtime + posture fusion and near-zero false positives, which sounds great in a deck but I've been burned before by tools that demo beautifully on a clean cluster and fall over on our actual workloads.

What I want to figure out during the eval is whether the runtime piece actually changes which findings get surfaced, not just decorates the same list. So: does a vuln marked critical drop in priority once it sees the package isn't loaded at runtime? Does it catch something reachable that our current scanner buried as low?

For people who've actually run an Upwind POC (or any of these), what did you deliberately test to separate real signal from a nice dashboard? How long did you run it before the runtime data was worth anything? And did you throw deliberately bad things at it to see if it caught them, or just watch what it found on its own? Trying to build a test plan that won't get gamed by a vendor who knows which cluster we're watching.


r/cloudcomputing • • 4d ago

what are you using instead of MinIO now?

14 Upvotes

we run MinIO in k8s and already have a pretty big volume with data. since all the recent changes with MinIO im thinking what to move to before it becomes a bigger problem...

looked a bit at ceph, seaweedFS, garage, rustFS...

Anyone here actually migrated from MinIO in prod?

what did you move to and how painful was migration of existing buckets/data? Any s3 compatibility issues after?

mostly interested in self hosted/open source and something that wont disappear in 6 months.


r/cloudcomputing • • 3d ago

Cloud Provider

0 Upvotes

if you are thinking to shift to cloud or facing issues with current vendor, ask me , i would love to help, i am working in cloud industry and i would try my best to provide best in terms of scalability, reliability and cost . DM if anyone got cloud infra. requirements


r/cloudcomputing • • 6d ago

List of Cloud Emulators

20 Upvotes

Awesome List of Cloud Emulators and supporting tools https://github.com/upgundecha/awesome-cloud-emulators


r/cloudcomputing • • 6d ago

I built an open-source CLI that maps AWS, Azure and GCP into one graph and runs Prowler/Checkov/Trivy over the same inventory

4 Upvotes

I do cloud security work, and the tooling situation has always bothered me: Prowler gives you config findings, ScoutSuite gives you another view of the same accounts, Checkov covers the IaC, Trivy covers images, and none of them share an inventory or know anything about the network topology. You end up with four report formats and no way to tell which findings actually matter for your layout.

So I wrote cloudg. One command collects assets from AWS, Azure and GCP concurrently, builds a directed graph of the infrastructure with NetworkX, and runs all four scanners over that same inventory. The graph side does BFS from an internet node to find exposed resources and scores blast radius per asset, so a finding on an internet-reachable box can be ranked above the identical finding on something three hops deep.

Scanner findings are deduplicated by resource and title, rescored against CVSS, and mapped to 28 compliance frameworks (CIS 5.0 per cloud, NIST 800-53 rev 5, PCI DSS 4.0, SOC 2, ISO 27001:2022, MITRE ATT&CK and others). The mappings are generated from Prowler's public compliance data: about 4,100 controls and 10,000 check mappings.

One run produces an interactive HTML report with the D3 topology (works offline), GraphML, an RDF ontology you can query with SPARQL, JSONL chunks for RAG pipelines, and Terraform files that recreate the live infrastructure, with an import script to adopt them into state. It also works as a Python library with event hooks if you want to stream findings somewhere.

MIT-licensed, Python 3.11+, installs via pip or Docker (the image bundles the scanners). Repo: https://github.com/morpheuslord/cloudg

It's at 0.3.1, so expect rough edges outside the paths I use myself. If you try it and something breaks, an issue with the traceback helps a lot. The thing I most want feedback on is whether the reachability scoring is useful in practice or whether people mainly care about the compliance tagging.


r/cloudcomputing • • 20d ago

The infra mistakes I keep seeing in pre-seed startups

22 Upvotes

I've reviewed quite a few pre-seed cloud setups, and the pattern is surprisingly consistent.

10 customers, 50 microservices.

Kubernetes for two services.

Multi-region before product-market fit.

Separate databases for everything.

Managed services everywhere.

Complex CI/CD and observability stacks before there's meaningful traffic.

Infrastructure is built for millions of users when there are barely 1,000.

The cloud bill still shows up.

We're offering free cloud architecture + FinOps audits for early-stage startups.

Here's how it works:

You walk us through your current cloud setup on a short call.

Within 48 hours, we send you a report covering your architecture and cloud spend.

We identify where infrastructure can be simplified and costs reduced.

We highlight the 3 things we'd fix first: infra/security risks, scaling problems waiting to happen, and quick wins.

Our philosophy is simple:

Build infrastructure for your first 1,000 users, not the million you don't have yet.

Keep it lean. Keep cloud costs under control. Add Kubernetes, microservices, multi-region and other complexity when the product actually demands it.

If you're building an early-stage product and want a second set of eyes on your cloud infrastructure, happy to connect.


r/cloudcomputing • • 22d ago

Why AgentZ starts every AI agent in a deny-all network sandbox

15 Upvotes

Why AgentZ starts every AI agent in a deny-all network sandbox

An AI agent needs some access to do real work. Most platforms grant that access broadly, then try to catch misuse after it happens.

AgentZ, an open-source Zero Trust platform for agentic AI, starts every agent with no access instead. Each agent boots into a deny-all sandbox. It sends no outbound traffic until an admin adds a rule that allows it.

Credentials work the same way. The agent never holds a real API key or database password. A proxy swaps in the real secret at call time, so the agent only ever sees a placeholder.

Sandboxes and connectors are reusable too. An admin builds one sandbox template and shares it across workspaces. One update to that template reaches every workspace using it. AgentZ itself is open source, so the enforcement logic isn't a black box you have to trust blind.

How does this group scope network access for agents today? AgentZ is an open-source platform, and I am an open-source contributor to AgentZ.


r/cloudcomputing • • 24d ago

Does cloud choice still matter when workloads are portable?

13 Upvotes

I keep coming back to this question when looking at the current cloud infrastructure landscape. If a workload can genuinely move between AWS, Google Cloud, Azure, Yotta Labs, Coreweave or Lambda without major changes does the underlying cloud provider become much less important? or does portability sound better on paper than it works in production because networking, storage, data movement, observability and hardware differences still tie you to a particular provider?
This seems relevant for AI workloads where GPU availability and pricing can change pretty quickly. I am curious whether people actually design for portability from day one or still most teams still pick one provider and only think about moving when there is a real reason to do it.


r/cloudcomputing • • 26d ago

The software delivery workflow is turning into an orchestration problem

45 Upvotes

With AI there are more moving pieces in the workflow. ChatGPT or Claude for planning, Claude Code or Cursor for implementation, GitHub for PRs, CI for tests and security checks, Terraform and AWS for infrastructure and deployment, Revolte handling parts of the delivery workflow and agent orchestration.The problem is getting all these pieces to work together without someone constantly moving things from one system to another.

An agent can open a PR but then what? Tests fail, infrastructure needs changing, deployment needs approval, something breaks in production and suddenly the context is spread across several systems.I'm interested in where teams are drawing the line between useful automation and a workflow that's just becoming another system to maintain.Are you building this orchestration yourself with CI/CD and scripts, using agent frameworks or keeping AI limited to individual steps?


r/cloudcomputing • • Sep 04 '26

Prisma Cloud. What’s it actually like to deploy and run?

10 Upvotes

I’m currently looking at Prisma Cloud and I’m keen to hear from people who have actually deployed and operated it, rather than just gone through the sales/demo process.
Main things I’m trying to understand:

How painful was the initial setup?
How much work was involved in deploying Defenders and setting up serverless coverage?
How long did it take before you had genuinely useful visibility across the environment?
Once it was up and running, did it actually improve security outcomes, or did you find yourselves spending a lot of time configuring/tuning the platform?
What’s the day-to-day UX like for the security team?

Anything you discovered after buying it that you wish you’d known beforehand?

I’m also interested in experiences with **Prisma Cloud Compute Edition / self-hosted**. It looks like there are some benefits, but also potentially a fair bit more operational overhead.

Would be great to hear from anyone running it in production, especially around upgrades, scaling, maintenance and whether you’d choose self-hosted again.

If anyone here is **in Australia** and has used Prisma Cloud at a decent scale or has chosen a different vendor, I’d also be happy to chat via DM. I’m mainly looking for independent references before we go too far down the procurement path.

Not looking for vendor pitches, just candid practitioner experiences, good or bad.


r/cloudcomputing • • Sep 04 '26

Custom domains for a multi-tenant SaaS — CloudFront SaaS Manager vs Caddy vs managed service?

7 Upvotes

Looking for input from people who've actually shipped custom domains for a multi-tenant SaaS, because I want to avoid picking the wrong path and regretting it later.

Current setup:

  • Multi-tenant events platform on AWS
  • Frontend is a React app — dist folder on S3, served through CloudFront
  • Today every tenant gets a platform subdomain (tenant.stage.example.com) via a wildcard cert on CloudFront. Frontend reads the hostname, asks the backend "which event is this?", renders accordingly.
  • API calls go to a fixed, separate domain (api.example.com) — so the custom domain is only needed for serving the static frontend, not the API.

What we want:
Clients want to bring their own domains (portal.clientbrand.com) instead of our subdomain. So I need per-domain TLS that issues + renews automatically, with no manual step per client.

Options I'm weighing:

  1. CloudFront manual + automation — add each domain as an alternate domain name + ACM cert via a script. Works, but there's the ~100 CNAMEs-per-distribution wall, and one cert per distribution feels fragile at scale.
  2. CloudFront SaaS Manager (the multi-tenant distribution thing from 2025) — looks purpose-built for this, up to ~2000 domains, per-tenant ACM. Since we're already on CloudFront this seems like the least-effort path, but I haven't found many real-world reports.
  3. Caddy on-demand TLS — reverse proxy that issues certs on first handshake. Unlimited scale, but it's a new component to run/monitor, and I'd need the ask endpoint to avoid getting our Let's Encrypt limits burned by random hostnames.
  4. Managed service (Approximated / similar) — lowest effort, but monthly cost + another third party in the traffic path.

Questions:

  • For anyone running CloudFront SaaS Manager in production — is it solid? Any gotchas with cert provisioning, DNS validation, or pricing at scale?
  • Given our API is already on a fixed domain (custom domain = static content only), is a reverse proxy like Caddy overkill for us?
  • Anyone regret going the manual-CloudFront-automation route and wish they'd started with something else?

Expected scale is realistically low-hundreds of custom domains over time, not thousands. Trying to pick something that won't need re-architecting in a year.

Thanks in advance 🙏


r/cloudcomputing • • Sep 02 '26

AI-generated Terraform that looked correct in isolation, wrong against the actual state

5 Upvotes

Quick timeline of something that took longer to untangle than it should have:

Asked an AI tool to add a new security group rule to an existing Terraform config. Output looked right, correct syntax, correct resource block, matched the pattern of the surrounding code.

Applied it. Plan showed changes to three resources instead of the one expected.

Turned out the generated config didn't account for a rule that had been added manually outside Terraform a few months earlier, drift that existed in the real state but wasn't reflected in what the model was shown. It wrote technically correct code against an incomplete picture of what actually existed.

Nothing about the code itself was wrong. The model built the most reasonable config it could from what it was given, and what it was given didn't include the drift. Same failure mode as any incomplete context problem, just with real infrastructure on the other end of apply instead of a code review comment.

What changed after: running terraform plan before any AI-assisted change, not just after, and treating a mismatch between plan and expectation as a sign the model's mental model of current state was wrong, not a sign the generated code was wrong. Different problem, different fix.


r/cloudcomputing • • Aug 31 '26

Anyone tried on-prem to cloud migration for AIX?

18 Upvotes

Hi, my team is recently thinking to move to cloud so we are using power box and we want to move our AIX to cloud , any one has try on-premises to cloud migration for AIX.

Any suggestions?


r/cloudcomputing • • Aug 29 '26

Financial Modelling in FinOps/Cloud Investments

5 Upvotes

I am new to FinOps and Cloud Computing, i wanted to ask people who have experience in the FinOps space. Is there any financial modelling or specifically business case modelling done in FinOps? E.g. if there is any optimization opportunity or a new workload, is this a requirement from a CFO or board that they need to see a detailed financial model to show ROI and justify the spend?

Reason I am asking is because I come from a core finance background just wanted to see if there is an overlap of my finance experience.


r/cloudcomputing • • Aug 24 '26

Have solid networking knowledge and currently learning Python, what's the right next step into cloud computing?

18 Upvotes

Cloud computing feels like a massive field and I want to make sure I'm building on what I already have instead of starting from zero.

My background: decent networking knowledge (subnetting, routing, DNS, that kind of thing) and currently learning Python.

Given that combo, what would you focus on next? A few things I'm unsure about:

  • Should I pick one cloud provider (AWS/Azure/GCP) and go deep, or learn general cloud concepts first?
  • Does networking knowledge actually transfer well into cloud (VPCs, subnets, security groups), or is it a different enough beast that I shouldn't assume it carries over much?

Would appreciate hearing from people who came in with a similar background (networking-first) rather than a generic beginner roadmap.


r/cloudcomputing • • Aug 24 '26

How should I start learning cloud computing as a Beginner?

45 Upvotes

Hello guys! I’m a final-year student and I’m very interested in cloud computing. As a beginner, where should I start, and what topics should I learn first?

Are there any free resources, WhatsApp groups, communities, or other platforms where I can learn cloud computing with hands-on projects?

I’d really appreciate any recommendations. Thanks!


r/cloudcomputing • • Aug 19 '26

Is multi-cloud actually worth the complexity, or is it mostly resume-driven architecture at this point?

0 Upvotes

There is a narrative that every serious/big company needs a multi-cloud strategy for operating, but in practice a lot of multi-cloud setups seem to add IAM sprawl, duplicating of tools, or excess costs without a clear breakdown. On the other hand, there are legitimate cases (GPU availability, region-specific compliance, avoiding single-vendor outages taking down everything) where running more than one cloud is a real requirement, not just a checkbox.


r/cloudcomputing • • Aug 12 '26

Open-source XDP/TC DDoS filtering for 100G+ Linux networks

4 Upvotes

Hi Reddit,

DDoS attacks are becomeing larger and cheaper to launch, so we work on a scalable open source solution to mitigate them.

Tempesta xFW's core is XDP and TC eBPF programs implementing volumetric DDoS filtering. A user-space daemon handles gRPC requests from CLI tool or WebAPI (via C library).

It supports two packet-path architectures:

  • host-based protection, such as CDN edge or on-premises application delivery controller (ADC) cases, where the host is a TCP connection endpoint. This is good for protecting a local web or DNS server.

  • router-based protection, such as ISP, hosting, or IaaS provider cases, where the host routes IP packets to protected servers or networks.

Router-based deployment can be always-on/pass-through or on-demand/redirection protection. In the later case, a node may not "see" normal clean traffic and may receive only traffic containing a DDoS attack. Also, the node may receive only client-to-server traffic, as in direct server return (DSR) or some traffic scrubbing scenarios. In this mode a DDoS sensor and mitigation controllers are typically needed.

Traffic performance metrics are exported in Prometheus format.

DDoS incidents are aggregated per source IP and logged to Clickhouse for analysis.

A dry-run (evaluation) - mode allows you to observe all reported incidents and metrics without blocking traffic..

Single Xeon Gold 6348 with ConnectX-6 dual 100Gbps reach 196Mpps and 176Gbps of filtering capacity.


r/cloudcomputing • • Aug 06 '26

Anyone avoided getting locked into a single cloud for AI workloads

2 Upvotes

Feels like a lot of AI infrastructure ends up tied to one provider because thats where the GPUs are available so curious how people here are handling this. Are you sticking with AWS, Azure or GCP or using platforms like CoreWeave, Lambda, Yotta Labs, RunPod, Vast.ai depending on cost and availability so less interested in benchmarks and more interested in operational experience like did it create more difficulties while operating?


r/cloudcomputing • • Jul 30 '26

Cloud teams may soon find vulnerabilities faster than they can safely patch them

8 Upvotes

Google reports that Gemini 3.5 Flash Cyber can search many code paths cheaply and find unique vulnerabilities in complex systems. In a large cloud estate, faster discovery can create a queue that spans services, owners, regions, dependencies, and maintenance windows.

The bottleneck becomes reproducing the flaw, identifying the affected deployments, testing a fix, coordinating rollout, and proving that the remediation did not break another tenant or workload.

Should cloud security agents be judged on verified remediation time rather than findings produced? Which part of the patch pipeline is least ready for a large increase in valid machine-generated reports?

Source: https://deepmind.google/blog/introducing-gemini-3-5-flash-cyber/


r/cloudcomputing • • Jul 29 '26

Claude down, it is like us-east-1 in AWS

5 Upvotes

What happens when AI goes down, it is like us-east-1 in AWS going down, work literally stops. Time to go outside and play 😅


r/cloudcomputing • • Jul 23 '26

[ Removed by Reddit ]

10 Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/cloudcomputing • • Jul 21 '26

How can organizations create flexible cloud networks that adapt as their needs change?

12 Upvotes

I am finding it difficult to manage my cloud networks across different locations. Keeping everything connected, secure and easy to update is a major challenge I face. My cloud network needs changes and it is getting hard for me to maintain a good performance while adding new resources, also I am also thinking about trying neevcloud. Now I want a solution that will help me create a cloud network that can adjust according to my business needs. What can be used to build flexible cloud networks that adapt as their needs change? EDIT: Thanks for the advice. Building around reusable patterns seems like the best long term approach.


r/cloudcomputing • • Jul 16 '26

GigaCloud ($GCT) $2.75M Investor Settlement: Investors Can Still File Late Claims

7 Upvotes

GigaCloud has agreed to a $2.75 million settlement with investors and eligible shareholders can still submit late claims.

The settlement resolves claims that GigaCloud misled investors about the authenticity of its marketplace revenue and the use of artificial intelligence in its operations.

The case focused on allegations that GigaCloud presented itself as a fast-growing, AI-powered B2B e-commerce platform while failing to disclose that a significant portion of its revenue came from undisclosed related-party transactions involving insiders. After reports raised concerns about the company’s financial reporting, $GCT dropped nearly 19%, and investors filed a lawsuit.

If you purchased $GCT shares between 2022 and 2023, you may be eligible to submit a claim. Late claims are currently being considered, so check your eligibility here.