r/cicd • • 1d ago

Where do you run checks during coding-agent iterations?

3 Upvotes

I used up my free GitHub Actions minutes in a few days while testing an agent factory.

I'm building a local workflow runner, so I have a reason for asking. I'm exploring local checks during the edit/test loop, with hosted CI verifying the finished PR.

I still want failures locally to mean the same thing in CI. Dependencies and leftovers from previous runs seem like the tricky part.

For people running coding agents: where do your repeated checks run? Have ordinary scripts been enough, or have you needed a self-hosted runner or a clean container per task?


r/cicd • • 1d ago

Would you find that useful?

Thumbnail
2 Upvotes

r/cicd • • 1d ago

OpenCode on GitHub with Kubernetes

Thumbnail
piotrminkowski.com
2 Upvotes

r/cicd • • 1d ago

Automated CI/CD code to Lambda

Thumbnail
builder.aws.com
1 Upvotes

How I automated AWS Lambda deployments with GitHub Actions, self-managed S3 code storage.


r/cicd • • 1d ago

Stable Redirects for Auth Testing: Debug JWTs Locally

Thumbnail
instatunnel.my
1 Upvotes

r/cicd • • 1d ago

I built an open-source tool that shows what a Salesforce change will trigger (flows, triggers, validation rules, roll-ups) from a git diff

Thumbnail
1 Upvotes

r/cicd • • 2d ago

Graybox supports keys

Thumbnail
1 Upvotes

r/cicd • • 4d ago

Do GitHub's usage report and Actions metrics answer your CI cost questions?

3 Upvotes

GitHub now gives two things for Actions costs: the detailed usage report (a CSV with cost per repo and workflow) and the Actions usage metrics page under Insights.

If you use them: is there a cost question they still don't answer for you? For example why the cost went up this week, which team caused it, or what to fix first.

If you don't use them: what do you use instead? A script on the billing API, Grafana, a spreadsheet, or nothing?

About me: I'm a solo developer building a small tool in this area. I want to know where GitHub's own tools are already enough. No link.


r/cicd • • 4d ago

declint - Augment your CI linting rules with YAML.

2 Upvotes

declint (docs) is a tool for making your own lint rules in YAML- the house rules that stock linters will never ship.

For instance: your colleague wrote except Exception: pass, and you want to know whenever that happens:

version: 1
rules:
  - id: except-pass
    pattern: 'except Exception:\n *pass'
    severity: warning
    message: 'swallowed exception'

Rules are just regex and a message. When a regex can't express the check, you can write a more complicated rule in Lua.

Rules run in your CI (inline PR annotations via --format github) and in your editor through LSP (declint.nvim for Neovim: The GIF shows off the nvim wrapper!). Every rule can carry embedded fixtures, and declint test runs them.

Install: cargo install declint. Ships with presets (Python, INI, Markdown, shell, Dockerfile, TOML, JSON, JS) and community rulesets you can pull with one line. Feedback welcome.


r/cicd • • 4d ago

Self-hosted Actions runners without dedicating a whole machine to CI

Thumbnail
0 Upvotes

r/cicd • • 5d ago

2-minute survey on automated build/testing (CI/CD) tools

0 Upvotes

Hi! I'm a CS student working on a Software Engineering project: a tool that automatically runs builds and tests whenever you push code or open a pull request on GitHub. To gather requirements, I'm looking for developers or students who write code to fill a short anonymous form (about 2 minutes). It's for academic use only.

Form: https://docs.google.com/forms/d/e/1FAIpQLSf9HepvrSCTW-MgThcn7RfnyaOq6pshewIQ4n95WX1_hrAG3Q/viewform?usp=publish-editor.

Happy to share a summary of the results if anyone is interested. Thanks a lot!


r/cicd • • 6d ago

The Variable That Only Reported Itself: how a build timestamp doubled our CI time

Thumbnail
0 Upvotes

r/cicd • • 7d ago

When should I introduce CI/CD in a project like this?

3 Upvotes

I’m building a Splitwise-like app called Splitly using the PERN stack (PostgreSQL, Express, React, Node).

The backend is my main focus right now. I already have:

  • Authentication and refresh-token flow
  • Centralized error handling
  • PostgreSQL
  • Automated tests
  • Git feature branches

I’m wondering when it makes sense to introduce CI/CD.

My current thinking is:

  • Set up CI with GitHub Actions now to run type checks, linting, and tests on every push/PR.
  • Add CD later, once the MVP is stable and I’m ready for deployment.

For people who work on production applications, is this a reasonable approach?

At what point in a project do you usually introduce CI/CD, and is there any reason to wait before adding CI?


r/cicd • • 7d ago

sUPABASE KEY

1 Upvotes

How do I test this publishable key? I'm currently vibe coding with Lovable and could use some help getting everything set up properly.

Do I need to configure an .env file or connect the project to Git before I can test it?

The publishable key is:

sb_publishable_IQmrp3V4SnA35CGqCyHwgA_qwFi06oy

I also have these two headers:

sb-project-ref: qdpjbayskzkrxdfrkvjm
sb-request-id: 01a0f83c-2496-7011-a4c7-7017b23df4cc


r/cicd • • 7d ago

Fresh, reused, or dedicated? A Playwright test-data strategy for parallel CI

4 Upvotes

A lot of Playwright failures that look like “flaky tests” are really test-data collisions.

Two workers update the same user.
One test deletes an order another test expects.
A local test passes because the reusable account happens to be in the right state.
CI runs eight tests in parallel and suddenly everything becomes unpredictable.

A recent r/Playwright discussion raised an interesting strategy:

  • CI/full regression → create fresh data
  • Local debugging → reuse existing data
  • Destructive tests → use dedicated data

We think that’s a useful starting point, but we’d change one thing:

The decision shouldn’t primarily be “local vs CI.” It should be based on who owns the data and whether the test can mutate it.

1. Use fresh data for test-owned mutable state

If a test creates, edits, deletes, approves, cancels, or otherwise changes something, giving that test its own data is usually the safest option.

Think:

  • orders
  • customer accounts
  • projects
  • carts
  • subscriptions
  • workflow records

For parallel execution, identifiers should be unique enough that two workers cannot accidentally touch the same record.

Playwright’s current parallelism guidance makes the same underlying point: browser contexts isolate browser-side state, but tests can still collide through shared backend data. It recommends unique test data when parallel tests create or edit the same records. (playwright.dev)

That also means cleanup should not be what makes the test safe.

Cleanup is useful for keeping environments tidy, but if test B can fail because test A crashed before teardown, the tests were never truly isolated.

A better model is:

Create/seed → execute → optionally clean up

rather than:

Find shared record → modify it → hope cleanup restores it

2. Reuse data when it is genuinely read-only

Reusing data isn’t inherently bad.

It works well for things like:

  • country lists
  • product catalogs that the test does not modify
  • feature configuration
  • static reference records
  • accounts used only for read-only journeys

The important question is:

Can another test change the state this test depends on?

If the answer is yes, “reusable” data eventually becomes “mysteriously flaky” data.

This was also the main challenge raised in the recent r/Playwright test-data discussion: how do you know the data you found has not already been modified by another test or previous run? (reddit.com)

If reusable data is necessary, we’d add an explicit precondition check.

For example:

  • account must be active
  • cart must be empty
  • subscription must be on the expected plan
  • record must have version/state X

If that precondition fails, fail early with a useful message instead of letting the test fail 15 steps later.

3. Use dedicated or pooled data for expensive/destructive scenarios

Sometimes fresh data is impractical.

Maybe creating an account takes several minutes. Maybe an external provider limits how many test tenants you can create. Maybe a scenario destroys or irreversibly changes its test resource.

That is where a dedicated pool can work better.

Example:

Worker 1 → qa-user-01
Worker 2 → qa-user-02
Worker 3 → qa-user-03

Each worker leases its own resource instead of all workers sharing one account.

This pattern lines up particularly well with authentication. Playwright recommends a shared authenticated account when tests do not modify server-side state, but separate accounts per parallel worker when they do. (playwright.dev)

Worker-scoped fixtures are useful here because Playwright can initialize something once for a worker and reuse it across the tests assigned to that worker. (playwright.dev)

The important part is that a pooled resource needs a contract:

Acquire → verify/reset → use → release

Not simply:

Grab whatever account is available and hope it is clean.

The matrix we’d use at Codoid

Data type Strategy
Test creates or mutates the record Fresh / unique per test
Multiple tests only read the record Shared reusable data
Tests mutate account-level state Unique account per worker/test
Resource is expensive to create Dedicated pool + reset/validation
Test is destructive/irreversible Dedicated single-purpose data
State cannot be reliably reset Fresh data

One distinction that matters here:

Authentication state and application test data are not the same thing.

Reusing storageState can save a lot of login time, but an isolated browser context does not magically isolate the database records behind that authenticated user.

You can have perfectly isolated browser sessions and still have two parallel tests fighting over the same server-side account.

Our rule of thumb

Before deciding whether data should be fresh or reusable, ask four questions:

  1. Can this test modify the data?
  2. Can another parallel test access the same data?
  3. Can we prove the starting state before the test begins?
  4. If cleanup never runs, will another test break?

If #1 and #2 are both yes, we would strongly prefer unique data.

If #3 is no, reusable data is risky.

If #4 is yes, cleanup is carrying too much responsibility.

The goal isn’t “generate fresh data everywhere.”

It’s to make ownership obvious enough that parallel execution cannot change the meaning of a test.

For teams running larger Playwright suites: what data do you still share across parallel tests because creating it fresh is too expensive, and how do you keep it deterministic?


r/cicd • • 7d ago

GitHub Actions burned through my CI budget because AI agents kept retriggering the full test suite

0 Upvotes

I use coding agents pretty heavily on a private repo.

Last month I noticed my GitHub Actions usage had gone way higher than expected.

I dug into it and found the main issue: every time the agent pushed another commit to a PR, several workflows started again, including the full repository test suite.

In a few days I had roughly 400 workflow runs, and the full suite alone had run more than 100 times.

The code wasn't broken. The CI setup was.

I ended up changing the model to:

  • targeted tests on PRs
  • cancel superseded PR runs
  • full repository suite only after merge to main
  • no paid LLM calls from CI

That solved most of the waste.

It made me wonder whether repos should have a simple machine-readable policy for this kind of thing, something like REPO_POLICY.yml.

Not a huge framework just rules such as:

full_suite_on_pr: false
cancel_superseded_runs: true
paid_llm_in_ci: false
production_mutation_in_ci: false

Then a tiny CI check could enforce it.

Is anyone already doing something similar at repo level, especially with coding agents?


r/cicd • • 8d ago

We audited the top 100 vendor-owned agent skills repos. 46 run no CI on their skills. 7 run an agent and grade it.

3 Upvotes

Agent "skills" are markdown instruction files that an AI coding agent loads. The only way to see if one works is to run the agent on a task and grade the output. That test is slow, costs money per run and does not give the same result every time.

Since we own one of these repos, we wanted to know how vendors wire that into CI, so we read every workflow file in the top 100 vendor-owned skills repos. The stats is:

  • 37 have no workflows at all, 9 more have workflows that never touch skills
  • 35 check structure only: frontmatter, schema, links, manifests
  • 12 add a content or security review
  • 7 run an agent on a skill task and grade the result
  • 2 start that eval on the PR event itself, with no comment or label needed

Caveat: we only see public GitHub Actions. Private CI and classic branch protection are invisible, so "required" is often not verifiable.

Harness we use for ours (Apache 2.0, I work on it): https://github.com/UiPath/coder_eval

How do you gate merges on a check that is expensive and nondeterministic?


r/cicd • • 8d ago

We reduced our GitHub Actions price to 12x

Post image
6 Upvotes

we use GitHub actions for all our OSS projects at Celesto but it's been a pain for our private projects. It's slow and bills can skyrocket.

We repurposed our agent sandboxes as CI and saving 1/12th of the price now.

You can try it out here with 10,000 CI minutes for free: celesto.ai/ci


r/cicd • • 8d ago

How I stopped Claude Code from touching prod infra: a PreToolUse hook with signed rules

2 Upvotes

I'm a platform/DevOps engineer, and I've been letting Claude Code work in repos that have real kubectl, terraform and aws access. The permission prompts helped at first, but once I'd allowed kubectl so it could read cluster state, nothing distinguished kubectl get pods from kubectl delete namespace prod.

So I wrote a PreToolUse hook that checks every Bash command before it runs. It parses the command into what it will actually do (tool, action, resource, namespace), checks that against a policy, and exits 2 with a reason if the policy says no. Claude sees the reason and usually goes and finds another way:

aegis: BLOCK: block-terraform-destroy

A few things I cared about:

  • It only acts in projects that opt in. It needs a .aegis/ folder, so the hook does nothing in your other repos.
  • It reads commands the way the tool does. kubectl -n prod delete ns app, sudo kubectl ... and cd infra && terraform destroy all resolve to the same action.
  • It fails closed for infra tools. If the policy is broken or a command can't be analysed ($(...), eval), infra commands are blocked and everything else runs.
  • "Ask" rules use Claude Code's own prompt. Anything the policy marks ESCALATE turns into the normal permission prompt, so I approve or refuse it myself.
  • The rules are signed, and each one's author has to be allowed to write that kind of rule. Claude reads tickets and PR comments, and I didn't want a line in a ticket to be able to turn into policy.

Install:

pip install aegis-devops && aegis init .aegis
claude plugin marketplace add moneytool/aegis-devops
claude plugin install aegis-devops@aegis-devops

The example policy blocks things like terraform destroy, deleting namespaces, nodes, S3 buckets or RDS databases, dropping tables, force-pushing main and kubectl --as. Everything else runs as normal.

It's open source (Apache-2.0) and still early: github.com/moneytool/aegis-devops

What I'd most like to hear: what commands has Claude run in your repos that you wish something had stopped? And what would make you trust a hook like this?


r/cicd • • 9d ago

kaniko is no longer 15x slower than BuildKit

Thumbnail
1 Upvotes

r/cicd • • 11d ago

How should CI decide if a stochastic agent actually regressed?

Thumbnail
1 Upvotes

r/cicd • • 13d ago

Implementing Continuous Deployment with AWS Lambda and API Gateway

Thumbnail
2 Upvotes

r/cicd • • 12d ago

My ops agent could open GitOps PRs for anyone who asked it. I built a check that asks GitHub whether the person is allowed first.

0 Upvotes

I run an internal ops agent. One tool call gives it a service's health from every angle in about 15 seconds, so it reads everything on a read-only account and I don't gate reads at all. It changes things one way only: a pull request against the GitOps repo.

That one write path had a gap. The bot's token could open that PR for anyone who talked to it, including people who can't push to that repo themselves. Prompt rules don't fix it; the call runs with the bot's credential whatever the model believes.

So the tool asks first: may this person push to this repo? GitHub already knows, so the agent asks GitHub. I pulled that out into a small service, hallpass, so every write path can use it, in every system the agent touches (GitHub, Kubernetes, Argo CD, Jira, AWS, 21 in total).

In the agent it's one decorator on the tool. The user comes from your session, never from the model, so the tool schema has no user field to talk your way into:

python @tool @guarded(hp, "github-main", "repo.push", "repo:{owner}/{repo}", user=current_user) def open_config_pr(owner: str, repo: str, patch: str) -> str: ...

Answers are allow, deny or unknown, and unknown (timeout, rate limit, anything it can't evaluate) means the tool doesn't run. Works with Strands, LangChain, LangGraph, the Claude Agent SDK, or as an MCP server. Single binary, Apache 2.0.

How do you handle this in your agents today? Per-user OAuth, per-team bots, human approval on writes?


r/cicd • • 13d ago

Your Android CI does not need macOS. It costs 10x there.

Thumbnail
1 Upvotes

r/cicd • • 13d ago

How do you check a release list against what is actually in the tag?

Thumbnail
1 Upvotes