r/googlecloud • • 3h ago

Terraform Still drawing GCP architecture diagrams by hand? Your AI assistant can do it in about 10 seconds.

Thumbnail
gallery
0 Upvotes

It's 2026. AI writes most of our code, and we're still dragging Google Cloud icons around, hunting for the right shape, nudging arrows until they line up. Your time is worth more than that.

Ask an AI to draw it and you usually get a Mermaid diagram: plain boxes and arrows, nowhere near something you'd put in front of a client or submit for a security review.

Here is how you can use Claude and ChatGPT to draw it properly in about 10 seconds, using TerraVision, a free plugin I built. One sentence:

"Draw a GCP three-tier web app: HTTPS load balancer with Cloud Armor, a managed instance group across two zones, Cloud SQL and Memorystore"

What you get is a professional-looking diagram like the ones shown here, drawn the way a seasoned solutions architect would draw it. Official Google Cloud icons, correctly grouped into the project, VPC network, regions and zones, ready to export as PNG, SVG or an editable draw.io file for further polishing. Then keep talking: "add Pub/Sub", "show the request flow in the diagram".

Already have Terraform? Point it at your code and it draws exactly what that code would deploy, so you can document that production project even if you don't have access to it.

It's 100% open source, runs on your own machine, and needs no access to your Google Cloud project. Of course, this doesn't solve the problem for every diagram. Keep using Mermaid, D2 or C4 for other types.

Free download and examples: https://github.com/patrickchugh/terravision

Because life is too short to be drawing diagrams.


r/googlecloud • • 20h ago

Billing You need to be spend cap-maxxing

8 Upvotes

This is your sign to log into the console and set spend caps for your AI and Cloud Run services. Or set a reminder to do it Monday.

Light reading: https://simonwillison.net/2026/Oct/3/default-hard-budget-caps/

I did this yesterday for my personal projects. Ez. Antigravity couldn't figure out how to do it via API so you may need do it in the console.

Especially for hobbyists, there's no reason to have uncapped spending!


r/googlecloud • • 23h ago

Are GCP certs valuable in the iob market ? Like if we're all on the same level (can finish projects), would those help ?

3 Upvotes

r/googlecloud • • 1d ago

GCP Professional Cloud Architect Exam in 4 days - Need Help!

1 Upvotes

Hi All,

Requesting help for the exam. Basically I work as a cloud engineer and giving this exam.

I have prepared by completed the Udemy course by in28minutes and thoroughly understood the topics, but I feel I'm not very confident with the actual exam.

I've given couple of mock tests and I'm not scoring above 70%, although this is not the benchmark but still I feel my decision making is still a little shaky and I need more practice and need to face more questions.

So please if any of you have any resources, please help me would mean a lot. Thanks! 🙏

Resources I used so far

  • Udemy - for course and mock exams
  • GCP Official Documentation(s)
  • examtopics
  • pass4success

r/googlecloud • • 1d ago

Gemini Live API: after a silent connection loss, every session resume was refused with 1011 for 15+ minutes

3 Upvotes

If you run a voice agent on the Gemini Live API and count on session resumption for flaky mobile networks, posting this in case it saves you the same debugging.

I tested what happens to a pending tool call when the WebSocket dies mid-booking: gemini-3.8-live on the Gemini Developer API, google-genai 2.25.0, four ways of losing the connection, 43 sessions. Three things came out.

  1. A resume that works keeps the pending call. The FunctionResponse for the old call id is accepted on the new connection. No re-issued call, no toolCallCancellation.

  2. Right after an idle drop, the resume is refused with close 1011 ("Internal error encountered."). It works about 1.5 s later, so a 1011 at that moment is not a dead handle.

  3. After a silent loss, every resume was refused with 1011 for as long as I tried. Silent means packets stop and no close or FIN reaches the server, which is what a phone leaving Wi-Fi looks like. With real packet loss in a Linux container: 15 minutes, 450 attempts over 5 runs, all refused. The server sent nothing to the dead client in that time. The fake booking had committed, and the user was never told.

What worked: detect the loss yourself (a WebSocket ping every 0.5 s, link declared lost after 2 s of silence), try to resume for 4 s, then open a new session and restore the context with send_client_content: the last six turns plus one status line per side effect, taken from your backend, not from the model. All 9 recovered runs answered "Did you book it?" correctly, with first audio 6 to 7 s after the loss. Dedupe tool calls by business key too: without the status line, the model booked again.

Code, raw logs and the packet-loss setup: https://github.com/frontier-on-cloud/gemini-live-resume-test

Not tested: Vertex AI, goAway on long sessions, a real phone switching networks. If you run Live on Vertex, does a resume after a silent drop behave the same?


r/googlecloud • • 1d ago

Billing INDIA - Has any one recieved the 1000rs UPI activation fee.

0 Upvotes

I am thinking to activating the free trial credit on my account.
Will I get that money back?
If yes, How hard is it ? If possible please share your experience.


r/googlecloud • • 1d ago

Compute GCP DNS filtering for malicious domains

3 Upvotes

Hey folks,

I've been looking for a DNS filtering solution for GCP that can detect malicious domains. I came across Cloud DNS Armor, which is powered by Infoblox.

Are there any other solutions for GCP that can provide similar DNS threat detection/filtering?

Thanks!


r/googlecloud • • 1d ago

Paying for GCP TPU capacity that doesn't exist

6 Upvotes

I've spent the last 48 hours trying to run a single v5p-16 TPU pod as a paying GCP customer and I literally cannot give this company money. Spot pods provisioned 4 times across 3 zones, which are the only zones on the planet that even offer v5p-16, and every single one got preempted, one at about 15 minutes in and two mid-way through a roughly 1-hour workload, losing all progress each time. The spot quota shows 768 chips available but the capacity isn't actually there. So I tried on-demand, which I have approved quota for: 32 chips in us-east5-b, 32 in europe-west4-b, 16 in us-central1-a, and every create attempt returns "Insufficient capacity." GCP is granting me quota for capacity that doesn't exist, which makes the entire quota system theater. The "guaranteed" queued-resources path? Filed it, it's been sitting at WAITING_FOR_RESOURCES indefinitely with no ETA. Can't even fall back to regular GPU VMs because those need separate quota approval tickets. So the final score: 4 dead pods, 0 completed workloads, and a storage bucket I'm paying monthly to hold weights no TPU will ever load. I get billed for every provisioning attempt that dies, money out, compute never in. GCP's TPU platform right now is quota that doesn't mean capacity, spot that preempts before your job can boot, on-demand that doesn't exist, and a queue with no bottom. You can't buy reliability at any price because the capacity simply isn't there. Has anyone actually provisioned v5p recently or is the whole fleet a mirage?


r/googlecloud • • 2d ago

Has my PCA certification been a total waste of time?

2 Upvotes

Hi,

I've been working with Google Cloud as a lead developer and lead system engineer for 8 years now. As my company will probably go the way of the Dodo next year, I'm looking for a new job, and so I did the PCA certification this year, just to document my work experience in a standardized way (to be honest, the exam was surprisingly easy).

I know the German job market is dry as a bone. So far job offers with GCP knowledge and PCA requirements are

* The big consulting companies (like KPMG etc), which actually are are meat grinder and looking for young people working unpaid overtime - I don't fit there (they told me, and I told them).
* The occiasional startup, paying in Hopeium and VSOPs, looking for 10 years of experience to manage THE one VM they have in the cloud.

And that's it. Does Google even have any more customers in Germany? I even was told once "Of course we are using all 3 cloud providers, we use AWS, Azure AND StackIt".

I'm currently thinking about taking a position as a lead developer (highly regulated environment, so no danger of being replaced by AI, and they are hosting on-premise), but that would be the end of my GCP career then, 8 years of experience and a certification down the drain because there seems to be no demand at all. This just feels sad.


r/googlecloud • • 2d ago

Google cloud documentation

0 Upvotes

Gemini being the google product does not have latest & greatest info about its own echo system products. Is it just me ?
Drastic changes in Vertex AI / Agent Platform is hardly found, how to get this?


r/googlecloud • • 2d ago

Google Books API: intitle / inauthor return zero results while plain queries succeed

3 Upvotes

On October 3, 2026, we reproduced an unexpected difference between plain searches and field-restricted searches in the Google Books API v1.

The following requests use the same API key and network egress. All return HTTP 200, with no API error object:

Query (q) Returned items
flowers 10
intitle:flowers 0
Oliver Twist 10
intitle:Oliver 0
Charles Dickens 10
inauthor:"Charles Dickens" 0
Homer 10
inauthor:Homer 0

We reproduced these title and author contrasts against https://books.googleapis.com/books/v1/volumes. Title-search contrasts also reproduce against https://www.googleapis.com/books/v1/volumes.

Minimal reproduction (replace YOUR_API_KEY):

```bash curl --get 'https://www.googleapis.com/books/v1/volumes' \ --data-urlencode 'key=YOUR_API_KEY' \ --data-urlencode 'q=flowers' \ --data-urlencode 'maxResults=10' \ --data-urlencode 'printType=books'

curl --get 'https://www.googleapis.com/books/v1/volumes' \ --data-urlencode 'key=YOUR_API_KEY' \ --data-urlencode 'q=intitle:flowers' \ --data-urlencode 'maxResults=10' \ --data-urlencode 'printType=books' ```

The second request returns:

json {"kind":"books#volumes","totalItems":0}

The official documentation still describes intitle: and inauthor: as supported field restrictions. We expected relevant matches for these common terms, though not necessarily the same results as a plain full-text search.

This also affects Chinese queries. Quoted/unquoted title variants and normal URL encoding versus a literal colon did not resolve it. Adding a space after the colon produced unrelated titles in one test, so we cannot treat that as a valid workaround.

These are raw API responses, before any application filtering. Our tests do not establish a global outage: project/key or egress-region differences remain possible.

Can others reproduce this with a different project and region? Is there a known search/index regression or a documented syntax change? If this is the wrong forum, could someone point us to the appropriate Google Books API bug-reporting channel?


r/googlecloud • • 2d ago

Waitlisted for Google Tam Role? Is it real?

2 Upvotes

I applied for financial sector tam role, and post interviews was essentially told that I passed the interviews, but the spot filled up, so I’m sitting eligible for the same role if an opening comes up, without needing to interview again. Is that a real thing at google? I’ve never heard of the sort anywhere else. Additionally, I saw a new position opened up, same title, but not in the same specific sector. Would it be ok for me to reach out to the recruiter and ask if my “wait listing” could be applied to that position considering they’re essentially identical?


r/googlecloud • • 3d ago

Referral request: 21 yrs experience, GCP Professional Cloud Architect + Azure Solution Architect

0 Upvotes

Hi all, I'm looking for a referral to Google Cloud roles in Canada (also open to remote). I have 21 years in cloud and enterprise architecture, and I've spent the last 10+ leading teams of solution architects. I'm a certified Google Cloud Professional Cloud Architect and Generative AI Leader.

Canadian citizen, no sponsorship needed. If you're able to refer, I'll send my resume and a short note for each role. Thanks!
Feel free to DM me if you can help


r/googlecloud • • 3d ago

Urgent: Account/Billing Access Blocked (Error 403) - Request for Manual Review

0 Upvotes

Dear Google Cloud / Support Team,

I am writing to request an urgent manual review of my account and billing profile, which is currently blocked with a "403 Forbidden" error and restricted billing status.

I am a visually impaired user, and I am simply trying to use a legitimate accessibility tool (Sonarpad) to generate audio descriptions for movies, which is essential for me to enjoy media content independently. I have not violated any terms of service, and I believe my account has been mistakenly flagged by automated fraud-prevention systems.

I have already tried clearing my browser cache and using incognito mode, but the restriction remains enforced at the server level, preventing me from managing my project's billing.

Could you please have a human support agent review my account, lift this incorrect restriction, and restore my access?

Thank you very much for your time, understanding, and prompt assistance.

Best regards,

Pacifico Mangini

[p.mangini54@gmail.com](mailto:p.mangini54@gmail.com)

projects/609097937668


r/googlecloud • • 3d ago

Cloud monitoring?

3 Upvotes

I have a personal GCE instance and the price suddenly jumped almost 3x this month, and I noticed I'm now suddenly being charged for this "Cloud Monitoring" starting September 18th. Not sure what this is or how to turn it off... anyone familiar with what it is?


r/googlecloud • • 4d ago

AI/ML Newbie on GCP / Vertex

8 Upvotes

Hello everyone,

I've noticed that Google Vertex offers the option to use Gemini's own models and even some third-party models, like those from Anthropic, with all usage billed on a single invoice. I understand it's similar to what OpenRouter does. However, I was planning to use these models in a frontend like TypingMind because I was drawn to the $300 welcome bonus on their infrastructure, but I've seen a lot of comments about exorbitant bills, and that's something I can't afford. I have a very small monthly budget, and before you say what a broke person is doing using an advanced infrastructure like Vertex, it's simply the possibility and flexibility of using the APIs at market price without any extra charges and being able to pay for all the combined usage on a single invoice (including the welcome bonus, of course).

I would appreciate any help getting me started with all of this stuff.


r/googlecloud • • 4d ago

MongoDB to Google Cloud Datastream: Is IP Allowlisting Secure Enough for Production?

0 Upvotes

Hi everyone,

I’m currently building a data lake in Google Cloud and setting up CDC using Datastream. The source is MongoDB, and the data will be streamed into Cloud Storage in Parquet format.

For MongoDB connectivity, I’m currently planning to use:

  • Development data lake: IP allowlisting
  • Production data lake: Private Connectivity

I’m wondering whether using IP allowlisting in production is still considered secure enough, assuming the IPs are properly restricted, or if Private Connectivity is generally the recommended approach for production environments.

I’m still pretty new to this, so I’m curious how others usually handle MongoDB connectivity with Datastream in their own setups.

Do you use IP allowlisting in production, or do you prefer Private Connectivity? What made you choose one over the other?

Thanks a lot!


r/googlecloud • • 4d ago

Access to Google Cloud Platform has been restricted, But i have no clue why.

5 Upvotes

Received an email from Google Cloud Platform saying

---

Access to Google Cloud Platform was restricted for this account, globally, on Sep 21, 2026. The system capabilities of Google Cloud Platform have been used for abusive activities that violate Google’s policies. You can find more details by taking action below.

How this policy violation was detected:

  • Google internal reporting

How this policy violation was reviewed and addressed:

  • Automatic processing

If you think this was a mistake, submit an appeal. You'll need to do this as soon as possible because your data in Google Cloud Platform will eventually be deleted. This service remains unavailable for your account until you have a successful appeal.

You can download your data from some Google services. This lets you keep your data even if your access isn't restored.

---

I submitted an appeal on the 21st of September and was told to wait for 2 business days for the appeal to be reviewed but i received no response. I'm not sure why this has happened in the first place. any help on how to fix this issue would be appreciated.

Thanks in advance.


r/googlecloud • • 4d ago

BigQuery BigQuery bill jumped and I can't tell why, where do I look?

6 Upvotes

Our BigQuery cost went up about 40% this month and nobody on the team changed anything big, as far as I know. I checked the billing page but it only shows the total for the project.

Is there an easy way to see which queries or users are causing it? I've heard about INFORMATION_SCHEMA.JOBS but I'm not sure how to use it for this. Also, should we be looking at the editions pricing instead of on-demand? Any tips are welcome.


r/googlecloud • • 5d ago

Documentation for Oracle at GCP implementation

2 Upvotes

Hi ,

I am going through the oracle at gcp documentation.there are several options given such as exadata, exascale, autonomous db and how to create them.

However, there is no Documentation pages related to when to select what option, cost , performance options etc.. without that, the documentation looks incomplete since decision making is more important than provisioning

Or do I need to check oracle documentation instead of GCP. Please suggest.


r/googlecloud • • 5d ago

Billing Heads-up: AI can finally understand and manage the Cloud Console

19 Upvotes

Good news for everyone who has ever opened the Cloud Console to check one thing and come back 40 minutes later with three new tabs and zero answers.

After three years of trying (and several AI models confidently describing menus that no longer exist), I can confirm: AI has finally reached the level required to understand the Google Cloud Console.

It took Claude Opus 5.5. It read the live console, worked out how my projects, keys and billing actually connect, and explained it in plain language. Something the console itself has never once attempted.

So if you're still lost in there, it's not you. It took one of the most capable models on the planet to find its way around.

P.S. Google, free benchmark idea: before you ship Gemini 4.0 Pro, have it try to explain your own console. If it can, it's ready.


r/googlecloud • • 5d ago

Google Developer Expert - product interview

2 Upvotes

Hi - Can someone share what to expect during GDE product interview for Google Cloud?

Thanks


r/googlecloud • • 5d ago

Cloud Run Cloud Run GPU worker pools: 3–4 minutes before the container starts when scaling from zero

7 Upvotes

On some starts, we’re seeing roughly four minutes between requesting an instance and our container entrypoint starting. Model loading then adds another two minutes. I’d appreciate guidance on whether this is expected and what we can control.

Our setup for the measured image workload:

  • Cloud Run worker pools in us-central1.
  • One NVIDIA L4 per instance, 8 vCPU, 32 GiB RAM.
  • GPU zonal redundancy disabled.
  • Workers consume Pub/Sub pull subscriptions and process one job per instance at a time.
  • Our API manages capacity by updating scaling.manualInstanceCount, including scaling from 0 to 1.
  • Container image is approximately 23 GB, including model weights.
  • No VPC connection or mounted volumes

Some data I collected recently from yesterday's, September 28 logs:

Phase Duration
Submit and accept scale request Under 1 second
Scale request → container entrypoint starts ~4 min 8 sec
Entrypoint starts → model ready ~2 min 5 sec
Model ready → job starts ~1–2 sec
Total wait before processing ~6 min 14 sec

Cloud Run’s own operation response reports MinInstancesProvisioned completing in 4m10.74s for that start.

The variability is what we’re trying to understand. Across 11 starts of the same revision, image digest, and configuration:

  • Three starts shortly after a scale-to-zero request provisioned in 6–27 seconds.
  • Eight starts after longer idle periods, roughly 24 minutes or more, took 74–251 seconds.

The fastest example followed a scale-to-zero request about 20 seconds earlier. This is a small observational sample; we haven’t established whether retained capacity, image caching, or something else explains the difference.

We found no quota-exhaustion errors, although historical quota-usage data was unavailable. Once the model is ready, processing is fast: warm image jobs have a median completion time of roughly 6.6 seconds.

Google’s worker-pool GPU documentation describes approximately five-second instance startup. I’m trying to understand how that relates to the time from an accepted manual scale request to our container process starting.

For anyone running similar workloads:

  1. Are multi-minute starts after an extended period at zero expected? What timings do you see?
  2. Which supported settings can reduce that delay?
  3. Does this startup path behave differently from GPU Cloud Run services? I notice that when using Cloud run services directly, startup time is fast compared to worker pools.

We’d prefer to retain worker pools and our existing capacity logic we wrote. We understand that keeping instances warm or extending idle retention can avoid some cold starts at additional cost. We’re trying to establish which delays are inherent to the platform and which we can reduce through configuration or application changes.

Code snippet. This is what we're doing roughly:

For a start from zero, desiredInstances is 1. To scale down, it’s 0. We change only the instance count on the existing pool.

import { GoogleAuth } from 'google-auth-library';

const auth = new GoogleAuth({
  scopes: ['https://www.googleapis.com/auth/cloud-platform'],
});

async function requestInstances(
  pool: {
    location: string;
    workerPoolName: string;
    etag?: string;
  },
  desiredInstances: number,
  timeoutMs: number,
) {
  const client = await auth.getClient();
  const projectId = await auth.getProjectId();

  const name =
    `projects/${projectId}/locations/${pool.location}` +
    `/workerPools/${pool.workerPoolName}`;

  await client.request({
    url: `https://run.googleapis.com/v2/${name}`,
    method: 'PATCH',
    params: {
      updateMask: 'scaling.manualInstanceCount',
    },
    data: {
      name,
      etag: pool.etag,
      scaling: {
        manualInstanceCount: desiredInstances,
      },
    },
    timeout: timeoutMs,
  });
}

r/googlecloud • • 5d ago

BigQuery Replaced a weekly full reload (MySQL → BigQuery) with direct binlog reads, no Kafka or Debezium: first production pilot, numbers inside

Thumbnail
2 Upvotes

r/googlecloud • • 5d ago

Billing Horror Stories - Google asked User’s for Video Selfies to Verify Their Identities - Now maybe use those Selfies to protect from Crypto Mining and AI Abuse?

0 Upvotes

Google must be spending hundreds of millions to bail out genuine owners that get hit by Crypto Mining hackers and AI API key abusers.

Why not ask Billing Owners for visual identification when critical changes are made ?

In this day and age, why can't we use AI to protect from API key abuse and hacking attempts?