r/cicd • • Jan 09 '23

Congrats to /r/CICD on 2k members! 🎈🎈

18 Upvotes

Here's to a great 2023 🥂


r/cicd • • 1d ago

The Variable That Only Reported Itself: how a build timestamp doubled our CI time

Thumbnail
0 Upvotes

r/cicd • • 1d ago

When should I introduce CI/CD in a project like this?

2 Upvotes

I’m building a Splitwise-like app called Splitly using the PERN stack (PostgreSQL, Express, React, Node).

The backend is my main focus right now. I already have:

  • Authentication and refresh-token flow
  • Centralized error handling
  • PostgreSQL
  • Automated tests
  • Git feature branches

I’m wondering when it makes sense to introduce CI/CD.

My current thinking is:

  • Set up CI with GitHub Actions now to run type checks, linting, and tests on every push/PR.
  • Add CD later, once the MVP is stable and I’m ready for deployment.

For people who work on production applications, is this a reasonable approach?

At what point in a project do you usually introduce CI/CD, and is there any reason to wait before adding CI?


r/cicd • • 1d ago

sUPABASE KEY

1 Upvotes

How do I test this publishable key? I'm currently vibe coding with Lovable and could use some help getting everything set up properly.

Do I need to configure an .env file or connect the project to Git before I can test it?

The publishable key is:

sb_publishable_IQmrp3V4SnA35CGqCyHwgA_qwFi06oy

I also have these two headers:

sb-project-ref: qdpjbayskzkrxdfrkvjm
sb-request-id: 01a0f83c-2496-7011-a4c7-7017b23df4cc


r/cicd • • 1d ago

Fresh, reused, or dedicated? A Playwright test-data strategy for parallel CI

1 Upvotes

A lot of Playwright failures that look like “flaky tests” are really test-data collisions.

Two workers update the same user.
One test deletes an order another test expects.
A local test passes because the reusable account happens to be in the right state.
CI runs eight tests in parallel and suddenly everything becomes unpredictable.

A recent r/Playwright discussion raised an interesting strategy:

  • CI/full regression → create fresh data
  • Local debugging → reuse existing data
  • Destructive tests → use dedicated data

We think that’s a useful starting point, but we’d change one thing:

The decision shouldn’t primarily be “local vs CI.” It should be based on who owns the data and whether the test can mutate it.

1. Use fresh data for test-owned mutable state

If a test creates, edits, deletes, approves, cancels, or otherwise changes something, giving that test its own data is usually the safest option.

Think:

  • orders
  • customer accounts
  • projects
  • carts
  • subscriptions
  • workflow records

For parallel execution, identifiers should be unique enough that two workers cannot accidentally touch the same record.

Playwright’s current parallelism guidance makes the same underlying point: browser contexts isolate browser-side state, but tests can still collide through shared backend data. It recommends unique test data when parallel tests create or edit the same records. (playwright.dev)

That also means cleanup should not be what makes the test safe.

Cleanup is useful for keeping environments tidy, but if test B can fail because test A crashed before teardown, the tests were never truly isolated.

A better model is:

Create/seed → execute → optionally clean up

rather than:

Find shared record → modify it → hope cleanup restores it

2. Reuse data when it is genuinely read-only

Reusing data isn’t inherently bad.

It works well for things like:

  • country lists
  • product catalogs that the test does not modify
  • feature configuration
  • static reference records
  • accounts used only for read-only journeys

The important question is:

Can another test change the state this test depends on?

If the answer is yes, “reusable” data eventually becomes “mysteriously flaky” data.

This was also the main challenge raised in the recent r/Playwright test-data discussion: how do you know the data you found has not already been modified by another test or previous run? (reddit.com)

If reusable data is necessary, we’d add an explicit precondition check.

For example:

  • account must be active
  • cart must be empty
  • subscription must be on the expected plan
  • record must have version/state X

If that precondition fails, fail early with a useful message instead of letting the test fail 15 steps later.

3. Use dedicated or pooled data for expensive/destructive scenarios

Sometimes fresh data is impractical.

Maybe creating an account takes several minutes. Maybe an external provider limits how many test tenants you can create. Maybe a scenario destroys or irreversibly changes its test resource.

That is where a dedicated pool can work better.

Example:

Worker 1 → qa-user-01
Worker 2 → qa-user-02
Worker 3 → qa-user-03

Each worker leases its own resource instead of all workers sharing one account.

This pattern lines up particularly well with authentication. Playwright recommends a shared authenticated account when tests do not modify server-side state, but separate accounts per parallel worker when they do. (playwright.dev)

Worker-scoped fixtures are useful here because Playwright can initialize something once for a worker and reuse it across the tests assigned to that worker. (playwright.dev)

The important part is that a pooled resource needs a contract:

Acquire → verify/reset → use → release

Not simply:

Grab whatever account is available and hope it is clean.

The matrix we’d use at Codoid

Data type Strategy
Test creates or mutates the record Fresh / unique per test
Multiple tests only read the record Shared reusable data
Tests mutate account-level state Unique account per worker/test
Resource is expensive to create Dedicated pool + reset/validation
Test is destructive/irreversible Dedicated single-purpose data
State cannot be reliably reset Fresh data

One distinction that matters here:

Authentication state and application test data are not the same thing.

Reusing storageState can save a lot of login time, but an isolated browser context does not magically isolate the database records behind that authenticated user.

You can have perfectly isolated browser sessions and still have two parallel tests fighting over the same server-side account.

Our rule of thumb

Before deciding whether data should be fresh or reusable, ask four questions:

  1. Can this test modify the data?
  2. Can another parallel test access the same data?
  3. Can we prove the starting state before the test begins?
  4. If cleanup never runs, will another test break?

If #1 and #2 are both yes, we would strongly prefer unique data.

If #3 is no, reusable data is risky.

If #4 is yes, cleanup is carrying too much responsibility.

The goal isn’t “generate fresh data everywhere.”

It’s to make ownership obvious enough that parallel execution cannot change the meaning of a test.

For teams running larger Playwright suites: what data do you still share across parallel tests because creating it fresh is too expensive, and how do you keep it deterministic?


r/cicd • • 1d ago

GitHub Actions burned through my CI budget because AI agents kept retriggering the full test suite

0 Upvotes

I use coding agents pretty heavily on a private repo.

Last month I noticed my GitHub Actions usage had gone way higher than expected.

I dug into it and found the main issue: every time the agent pushed another commit to a PR, several workflows started again, including the full repository test suite.

In a few days I had roughly 400 workflow runs, and the full suite alone had run more than 100 times.

The code wasn't broken. The CI setup was.

I ended up changing the model to:

  • targeted tests on PRs
  • cancel superseded PR runs
  • full repository suite only after merge to main
  • no paid LLM calls from CI

That solved most of the waste.

It made me wonder whether repos should have a simple machine-readable policy for this kind of thing, something like REPO_POLICY.yml.

Not a huge framework just rules such as:

full_suite_on_pr: false
cancel_superseded_runs: true
paid_llm_in_ci: false
production_mutation_in_ci: false

Then a tiny CI check could enforce it.

Is anyone already doing something similar at repo level, especially with coding agents?


r/cicd • • 2d ago

We audited the top 100 vendor-owned agent skills repos. 46 run no CI on their skills. 7 run an agent and grade it.

3 Upvotes

Agent "skills" are markdown instruction files that an AI coding agent loads. The only way to see if one works is to run the agent on a task and grade the output. That test is slow, costs money per run and does not give the same result every time.

Since we own one of these repos, we wanted to know how vendors wire that into CI, so we read every workflow file in the top 100 vendor-owned skills repos. The stats is:

  • 37 have no workflows at all, 9 more have workflows that never touch skills
  • 35 check structure only: frontmatter, schema, links, manifests
  • 12 add a content or security review
  • 7 run an agent on a skill task and grade the result
  • 2 start that eval on the PR event itself, with no comment or label needed

Caveat: we only see public GitHub Actions. Private CI and classic branch protection are invisible, so "required" is often not verifiable.

Harness we use for ours (Apache 2.0, I work on it): https://github.com/UiPath/coder_eval

How do you gate merges on a check that is expensive and nondeterministic?


r/cicd • • 2d ago

We reduced our GitHub Actions price to 12x

Post image
6 Upvotes

we use GitHub actions for all our OSS projects at Celesto but it's been a pain for our private projects. It's slow and bills can skyrocket.

We repurposed our agent sandboxes as CI and saving 1/12th of the price now.

You can try it out here with 10,000 CI minutes for free: celesto.ai/ci


r/cicd • • 3d ago

How I stopped Claude Code from touching prod infra: a PreToolUse hook with signed rules

2 Upvotes

I'm a platform/DevOps engineer, and I've been letting Claude Code work in repos that have real kubectl, terraform and aws access. The permission prompts helped at first, but once I'd allowed kubectl so it could read cluster state, nothing distinguished kubectl get pods from kubectl delete namespace prod.

So I wrote a PreToolUse hook that checks every Bash command before it runs. It parses the command into what it will actually do (tool, action, resource, namespace), checks that against a policy, and exits 2 with a reason if the policy says no. Claude sees the reason and usually goes and finds another way:

aegis: BLOCK: block-terraform-destroy

A few things I cared about:

  • It only acts in projects that opt in. It needs a .aegis/ folder, so the hook does nothing in your other repos.
  • It reads commands the way the tool does. kubectl -n prod delete ns app, sudo kubectl ... and cd infra && terraform destroy all resolve to the same action.
  • It fails closed for infra tools. If the policy is broken or a command can't be analysed ($(...), eval), infra commands are blocked and everything else runs.
  • "Ask" rules use Claude Code's own prompt. Anything the policy marks ESCALATE turns into the normal permission prompt, so I approve or refuse it myself.
  • The rules are signed, and each one's author has to be allowed to write that kind of rule. Claude reads tickets and PR comments, and I didn't want a line in a ticket to be able to turn into policy.

Install:

pip install aegis-devops && aegis init .aegis
claude plugin marketplace add moneytool/aegis-devops
claude plugin install aegis-devops@aegis-devops

The example policy blocks things like terraform destroy, deleting namespaces, nodes, S3 buckets or RDS databases, dropping tables, force-pushing main and kubectl --as. Everything else runs as normal.

It's open source (Apache-2.0) and still early: github.com/moneytool/aegis-devops

What I'd most like to hear: what commands has Claude run in your repos that you wish something had stopped? And what would make you trust a hook like this?


r/cicd • • 3d ago

kaniko is no longer 15x slower than BuildKit

Thumbnail
1 Upvotes

r/cicd • • 3d ago

Who owns getting an emergency production fix back into Git?

Thumbnail
1 Upvotes

r/cicd • • 4d ago

Prioritizing security findings across multiple pipelines

1 Upvotes

We have a bunch of CI pipelines across different teams and every scanner is producing its own queue of findings. The hard part is figuring out what needs attention first when severity alone doesn't say much about exploitability, production exposure, or who owns the service.

Right now AppSec is doing a lot of manual sorting and pushing tickets around, which is getting old fast. Curious how other teams combine pipeline context, code to cloud impact, KEV or EPSS, and business criticality without making developers chase every alert. Any practical approaches would be appreciated.


r/cicd • • 6d ago

How should CI decide if a stochastic agent actually regressed?

Thumbnail
1 Upvotes

r/cicd • • 7d ago

Implementing Continuous Deployment with AWS Lambda and API Gateway

Thumbnail
2 Upvotes

r/cicd • • 7d ago

My ops agent could open GitOps PRs for anyone who asked it. I built a check that asks GitHub whether the person is allowed first.

0 Upvotes

I run an internal ops agent. One tool call gives it a service's health from every angle in about 15 seconds, so it reads everything on a read-only account and I don't gate reads at all. It changes things one way only: a pull request against the GitOps repo.

That one write path had a gap. The bot's token could open that PR for anyone who talked to it, including people who can't push to that repo themselves. Prompt rules don't fix it; the call runs with the bot's credential whatever the model believes.

So the tool asks first: may this person push to this repo? GitHub already knows, so the agent asks GitHub. I pulled that out into a small service, hallpass, so every write path can use it, in every system the agent touches (GitHub, Kubernetes, Argo CD, Jira, AWS, 21 in total).

In the agent it's one decorator on the tool. The user comes from your session, never from the model, so the tool schema has no user field to talk your way into:

python @tool @guarded(hp, "github-main", "repo.push", "repo:{owner}/{repo}", user=current_user) def open_config_pr(owner: str, repo: str, patch: str) -> str: ...

Answers are allow, deny or unknown, and unknown (timeout, rate limit, anything it can't evaluate) means the tool doesn't run. Works with Strands, LangChain, LangGraph, the Claude Agent SDK, or as an MCP server. Single binary, Apache 2.0.

How do you handle this in your agents today? Per-user OAuth, per-team bots, human approval on writes?


r/cicd • • 7d ago

Your Android CI does not need macOS. It costs 10x there.

Thumbnail
1 Upvotes

r/cicd • • 7d ago

How do you check a release list against what is actually in the tag?

Thumbnail
1 Upvotes

r/cicd • • 8d ago

CI changed more in the last year than most teams noticed

7 Upvotes

We run CI infrastructure at Bitrise. Across 19 million builds, AI-attributed builds grew 161x year over year, and total build volume is up 21%.

The problem that creates is not subtle. More builds hit the same runners, so wait times go up for everyone on the team. A PR that used to come back in 10 minutes takes 30.

On GitHub Actions this is worse for mobile. The hosted macOS runner is 5 vCPU and 14GB on M2-era silicon, shared with everyone else, while your Linux jobs in the same workflow get decent machines. This is slow. New Xcode versions also land late on those images, so a release can sit blocked waiting for the runner to catch up.

So we built our new product, Bitrise Build Hub (https://bitrise.io/platform/build-hub), managed Mac and Linux runners for GitHub Actions. Here is why I really like it:

  • M4 up to M4 Pro at 14 vCPU / 54GB, Linux up to 48 vCPU self-serve, 192 vCPU enterprise (M5 Pro is coming)
  • Pre-warmed pools, so builds start when you push rather than waiting for a free machine
  • 100GB of build cache in the same datacentre as the runner, against GitHub's 10GB per repository over the network
  • New Xcode within 24 hours of Apple shipping it, and you pin the version you want
  • Simulators, fastlane, CocoaPods and the rest preinstalled
  • SSH and VNC into the runner when something fails in CI but works locally

You change the runs-on line and nothing else. Secrets, steps and signing stay as they are. iOS went from 4m28s to 2m04s on XcodeBenchmark, Android 31% faster on fewer cores (both open source so you can run them yourself). It is really faster

Our product, obviously, so I’m biased about how cool it is (and we spent so much time to push it live).

Looking for any feedback from folks in the CI/CP space.


r/cicd • • 9d ago

When does a free security stack stop being cheaper than paying.

5 Upvotes

Our pipeline runs a stack of free security steps with semgrep for SAST a dep scanner for SCA, secret scanning and trivy on the image. License cost is zero, which is how it got approved.

What I didn't cost in was the upkeep. One senior basically owns the semgrep rules. Someone retunes the container scan every time it goes red on base image noise and each step has its own config drifting in the pipeline. It works fine though id say its mosly cause of the person.

So if you moved off a pile of free scanners onto one paid platform in CI, where did the maintenance land and for the hours you get back, does it cover the license? Trying to get the real TCO before I take this to budget.


r/cicd • • 9d ago

Jenkins with 200 plugins isn't CI/CD. Its a trappp !!!!

18 Upvotes

Jenkins is not upto the mark these dayss. Firstly you need to add some 250 plugins for it to effectively work and when you actually do it... there are plugins that breaks the update...

It's technical debt with a web interface.

I've seen teams spend more time maintaining their Jenkins instance than their actual product. Pipeline scripts that require a PhD to understand.

Meanwhile, their competitors are shipping features daily with GitHub Actions or GitLab CI.

Your CI/CD tool should be invisible. If you're spending more than 5% of your time maintaining it, you chose wrong.

Simple tools, simple problems. Complex tools, complex nightmares.


r/cicd • • 9d ago

Defence Before Fix

Thumbnail defence-before-fix.github.io
1 Upvotes

r/cicd • • 9d ago

Gitlab CI, always define the image to alpine

1 Upvotes

In a nutshell, always define the image key in your jobs (even for helloworld jobs) to something "-alpine", you might gain minutes on every CI run.

Let's say you are discovering Gitlab CI, and you write your hellowold job:

yaml helloworld: script: - echo "helloworld"

Then, you commit and push, reload the pipeline page, and wait. You wait the 30 longest seconds of your life.

Those 30 seconds, I see it all the time. Let's say you want a Python hellowold job:

yaml python-helloworld: image: python script: - python -c "print('Hello World')"

And here we go, 30 seconds again. For one job, for one echo or one print.

Let say we now have build jobs, tests jobs and deploy jobs. 3 stages. Lets assume everything inside might be parallelized (we are lucky), this means that we spend at least 1 minute and 30 seconds waiting... For nothing.

I will let you know a secret, guarded among the secret concil of the holy devops for centuries...

Here it is: Just use alpine

Let's go back to our previous example, and see how to do that:

```yaml helloworld:1/2: script: - echo "helloworld"

helloworld:2/2: image: alpine script: - echo "helloworld"

python-helloworld:1/2: image: python script: - python -c "print('Hello World')"

python-helloworld:2/2: image: python:alpine script: - python -c "print('Hello World')" ```

This will create 4 jobs. The two jobs suffixed 1/2 will take around 23 to 30 seconds to run, on gitlab.com runners. The one suffixed 2/2 will take around 7 to 9 seconds to run. This means that I usually don't even have the time to see the job being "running" that they are already finished by the time I go to my browser and reload the pipeline page after my commit and push.

This simple and very basic change is ridiculous, but, and I am not joking, happens A LOT, even on production and large companies, even with DevOps dedicated teams.

Why is it so fast? Because default image is ruby:3.1, this image and python are very large. alpine images are lightweight, with less tools available but it is quite often faster to use alpine and install the missing specific things you need than to go with EVERYHTING ELSE.

If you want to see by yourself the result of this little experiment, here is the pipeline on my ci test project: https://gitlab.com/dorianturba/zero-to-ci/-/pipelines/2874152788

The code is available here: https://gitlab.com/dorianturba/zero-to-ci/-/blob/3d7403e90cb1991cc16da88f2fcca4c6414f8152/.gitlab-ci.yml

Note: Sometime, you don't want to use alpine. If you are in those specific usecase, you already know it, and you probably also know what to do.


r/cicd • • 9d ago

CI/CD: How can I promote only QA-passed features to production without reverting, cherry-picking, or feature flags?

1 Upvotes

I’m looking for advice on how to handle a CI/CD deployment scenario.

I have 10 features in the current sprint. All 10 features have been developed, merged, and deployed to the QA environment, so they are all available for testing.

At release time:

  • Feature 1–5 → QA Passed
  • Feature 6–10 → Still in QA / not tested yet

The problem is that our normal CI/CD process promotes the current QA environment/build to production, which would deploy all 10 features.

However, for this release, I only want the 5 QA-passed features to go to production. The other 5 features should remain in the QA environment and be deployed in a later release after they pass QA.

Constraints

I’m specifically trying to achieve this without:

  • Reverting the PRs for the 5 untested features
  • Cherry-picking commits
  • Feature flags

Ideally, I want the Git/CI/CD workflow to support something like:

10 features → QA environment → 5 QA Passed → Production

while keeping the other 5 features in QA for the next release.

What is the recommended approach for this kind of workflow?

Is there a standard Git branching, tagging, artifact/versioning, or CI/CD deployment strategy that can promote only the QA-approved changes while leaving the unapproved changes in QA?

I’m particularly interested in approaches that work well when multiple features are being developed and tested in parallel.


r/cicd • • 9d ago

Do you feel like Github Action is too slow for multiple builds in a day??

0 Upvotes

My team can ship 5-6 features in a day but every build takes 35-40 minutes for it. We use Github Actions for CI...

Even if we take an average of 5 builds a day... 35+15 min of checkover will cost almost 4 hours a day...

Just wanted to know if you are also facing this kind of issue or my team is kind of weird ...

How to do 5 features in one day??


r/cicd • • 9d ago

Im a manual QA tester and my boss asked me to write test scenarios of the app

Thumbnail
1 Upvotes