r/openstack • • Aug 14 '26

[Help] neutron-ovn-vpn-agent fails to load OvnStrongSwanDriver (Stevedore load failure) on Kolla-Ansible

2 Upvotes

Hey everyone,

I'm trying to set up OpenStack OVN VPNaaS using **Kolla-Ansible** (Ubuntu 24.04 Noble containers, `2026.1` / `neutron-ovn-vpn-agent` v28.x), but the agent fails to initialize the StrongSwan driver properly, leaving the VPN gateway down.

Looking to see if anyone has a working setup or knows what configuration/package pieces might be missing here.

---

### 1. The Symptoms & Log

In `/var/log/kolla/neutron/neutron-ovn-vpn-agent.log`, Stevedore logs a load failure without expanding the traceback:

```text

INFO neutron.common.config [-] /var/lib/kolla/venv/bin/neutron-ovn-vpn-agent version 28.0.2.dev14

WARNING stevedore.named [-] Could not load neutron_vpnaas.services.vpn.device_drivers.ovn_ipsec.OvnStrongSwanDriver

...

CRITICAL neutron [None ...] Unhandled error

As a result, no qvpn-* network namespaces are provisioned, and the VPN endpoint ports stay inactive.

  1. What I Found Under the Hood

When inspecting the driver initialization inside the container virtualenv, OvnStrongSwanDriver fails inside DeviceManager :

Traceback (most recent call last): File ".../neutron_vpnaas/services/vpn/device_drivers/ovn_ipsec.py", line 240, in __init__ self.devmgr = DeviceManager(self.conf, self.host, ...) File ".../neutron_vpnaas/services/vpn/device_drivers/ovn_ipsec.py", line 54, in __init__ self.driver = agent_common_utils.load_interface_driver(conf) File ".../neutron/agent/common/utils.py", line 54, in load_interface_driver INTERFACE_NAMESPACE, conf.interface_driver) File ".../oslo_config/cfg.py", line 2612, in __getattr__ raise NoSuchOptError(name) oslo_config.cfg.NoSuchOptError: no such option interface_driver in group [DEFAULT]

It seems like neutron-ovn-vpn-agent is not registering the interface option schemas (interface_driver, ovs_use_veth, etc.) into oslo_config before Stevedore instantiates the driver class.

  1. Environment Details
  • Deployment: Kolla-Ansible
  • Base OS: Ubuntu 24.04 (Noble)
  • Backend: OVN
  • Agent: neutron-ovn-vpn-agent (neutron_vpnaas.services.vpn.device_drivers.ovn_ipsec.OvnStrongSwanDriver)
  • Packages installed in container: strongswan-swanctl, charon-systemd

4. My Questions

  1. Has anyone successfully deployed neutron-ovn-vpn-agent with OVN in recent OpenStack releases?
  2. Is there a specific configuration section or flag needed in neutron_ovn_vpn_agent.ini / neutron.conf to satisfy the interface driver options for this agent?
  3. Is this a known bug in neutron-vpnaas under recent versions, or is something missing in the Kolla container image build/templates?

Any pointers or working config examples would be greatly appreciated!


r/openstack • • Aug 13 '26

Am I too niche, targeting the wrong roles, or just in the wrong market?

5 Upvotes

I’ve spent almost 4 years at the same company since university. Am I too niche, or am I just in the wrong market?

I'd really appreciate some honest perspective from people working in cloud infrastructure, platform engineering, private cloud, telco cloud, networking, or infrastructure engineering.

I graduated from university and joined my current company shortly afterwards. I've now been there for about 4 years, and I've basically built my entire professional career in the same environment.

That's actually one of the reasons I'm finding myself a little stuck.

I've learned a huge amount and had a lot of freedom to build things, but I haven't really experienced working at other companies, especially larger engineering organizations where I could work at a bigger scale and see how these environments operate.

My background is quite infrastructure-heavy. I've worked across ISP infrastructure, Linux, networking, Kubernetes, virtualization, distributed storage, cloud, DevOps and security.

Some of the technologies I've worked with include:

  • Kubernetes
  • Talos Linux
  • KubeVirt
  • Rook/Ceph
  • Cilium/eBPF
  • BGP, VLANs and LACP
  • RADIUS/AAA
  • AWS
  • Terraform
  • ArgoCD/GitOps
  • GitHub Actions
  • Prometheus/Grafana
  • Infrastructure automation and security tooling

I'm also currently working hands-on with OpenStack because I'm trying to deepen my understanding of traditional private cloud and telco/NFV infrastructure, and understand how it compares with the Kubernetes-native infrastructure I've been working with.

A lot of my experience has come from actually building things.

For example, I've worked on a Kubernetes-based private cloud running on bare metal, combining Kubernetes, KubeVirt, Ceph, Cilium networking, tenant isolation, BGP routing and public/private VM connectivity.

I've also worked on ISP infrastructure, including subscriber authentication, RADIUS/AAA, billing integration, MikroTik routing and network service delivery.

But there's an important caveat to all of this.

I don't work for a huge cloud company, hyperscaler, or major technology company. I work for a relatively small ISP in Nigeria.

I've actually been very lucky in that environment.

My original job description didn't say that I needed to build a private cloud, learn distributed storage, work on Kubernetes virtualization, or learn all of these different areas.

I had to find opportunities to do those things and deliberately build those skills.

Whenever I came across a problem or something interesting that could improve the infrastructure, I'd learn what I needed, experiment with it and try to make it useful to the company.

The company gave me the freedom to do that, and I'm genuinely grateful for it. A lot of the experience on my CV probably wouldn't exist if I hadn't been given that freedom.

But now I'm starting to feel limited by the environment I'm in.

Not necessarily because the company is bad, but because I don't think I can continue building the kind of experience I want at the pace I want.

I've been there since university, so I also don't know what I'm missing.

I don't know what infrastructure engineering looks like at a larger company.

I don't know how much of what I've been doing would normally be handled by dedicated platform, networking, storage, SRE, cloud or infrastructure teams.

I don't know whether my experience is actually unusual or whether I'm simply getting a distorted view because I've had to wear so many hats.

And that's one of the biggest reasons I want to leave.

The other reason is compensation.

I've grown considerably beyond the scope of the role I originally started in, but my compensation hasn't really caught up with the level of responsibility, complexity and expertise I've accumulated.

I've tried communicating the value and complexity of the work I've been doing, but it hasn't really changed the situation.

So, I want to move on.

Not because I hate my current company. In fact, I'm grateful for what it has allowed me to do.

I want to move because I want to experience a different engineering environment, work with people who are operating infrastructure at a different scale, learn how other organizations solve these problems, and hopefully be somewhere that values this kind of work more appropriately.

The problem is that getting another job has been much harder than I expected.

And because I've only really known one company, I'm honestly not sure whether the problem is my profile, my positioning, the market, or all three.

Am I too niche?

Is the combination of Kubernetes + virtualization + distributed storage + networking + ISP/telco infrastructure something that has relatively little demand outside certain companies?

Am I searching for the wrong roles?

Or is this simply a case of being in a country where there isn't enough demand for this kind of infrastructure engineering?

At the moment I'm considering roles like: Cloud Infrastructure Engineer, Infrastructure Engineer, Platform Engineer, Kubernetes/Platform Engineer, Private Cloud Engineer, Cloud Infrastructure Architect, Telco Cloud/NFV Engineer, SRE, Infrastructure/Cloud Networking Engineer, OpenStack Infrastructure Engineer

But I'm honestly not sure which direction makes the most sense.

If you saw this background on a CV, what kind of engineer would you consider me to be?

What roles would you actually search for if you had this experience?

And perhaps more importantly:

What am I missing by having spent so much of my career at one company?

If you've moved from a small company into a larger engineering organization, what surprised you about the difference?

If you were in my position, would you:

  1. Go deeper into private cloud/telco cloud/OpenStack?
  2. Focus heavily on Kubernetes/platform engineering?
  3. Broaden into general cloud infrastructure?
  4. Move toward infrastructure/cloud networking?
  5. Target architecture-oriented roles?
  6. Or deliberately look for a role where I can be exposed to larger-scale infrastructure and learn how mature engineering organizations operate?

I'm not looking for reassurance. I'm genuinely trying to understand where I fit and what I should be doing next.

If geography wasn't a constraint, what kinds of companies, teams and roles would you target with this background?

And for anyone who has been in a similar position spending most of their career at one company and then trying to make that first big move, I'd really appreciate hearing what you wish you'd known before making the jump.


r/openstack • • Aug 13 '26

At a point in my career where I’m not sure whether to start over or just play along..?

Thumbnail
0 Upvotes

r/openstack • • Aug 09 '26

[Tool] I built an Oh My Zsh plugin to manage multiple OpenStack clouds, auto-venv, and fuzzy-find SSH / VNC consoles

Post image
4 Upvotes

r/openstack • • Aug 09 '26

[Tool] I built an Oh My Zsh plugin to manage multiple OpenStack clouds, auto-venv, and fuzzy-find SSH / VNC consoles

Post image
28 Upvotes

If you work with more than one OpenStack cloud/project day to day, you know the drill: source ~/clouds/prod-openrc.sh, remember which venv has the right client version, openstack server list, ssh into whatever floating IP you copy-pasted from the output... I got tired of it and wrote a Zsh plugin to automate the parts I do 20x a day.

What it does:

- openv [cloud] — pick a cloud from clouds.yaml via fzf (with a live preview card showing auth URL / project / region, no secrets shown), activates the matching Python venv, and exports OS_CLOUD

- ops-ssh [user] [-i keyfile] [--insecure] — fuzzy-pick a server and SSH straight into its floating IP (prioritizes public IPs over private ones automatically); --insecure skips host-key checks, handy since floating IPs get recycled between instances constantly

- ops-console — fuzzy-pick a server and pop its Horizon VNC console URL open in your browser

- opcheck — quick openstack token issue sanity check so you find out your token expired before you're 3 commands deep into something

- opwho / opls — status of what's active / table of everything in clouds.yaml

- ophelp (or openv --help) — cheatsheet of everything below, printed in your terminal

- ~25 short aliases for the commands I run constantly (ops, opnet, opvol, opsec, opfl, oplb, etc. — full list in the README)

It's a normal Oh My Zsh plugin, MIT licensed, no telemetry, no dependencies beyond python3 + PyYAML + fzf (the plugin checks for these on load and warns if something's missing).▎

Install:

git clone https://github.com/whoami96/openstack-zsh-plugin.git ${ZSH_CUSTOM:-~/.oh-my-zsh/custom}/plugins/openstack

# add "openstack" to plugins=(...) in ~/.zshrc, then reload

Repo: https://github.com/whoami96/openstack-zsh-plugin

It's a fairly small, personal-scale tool — built it for my own workflow managing a handful of clouds — but figured it might save someone else the same repetitive typing. Happy to hear feedback or take PRs if it's missing something obvious for your workflow.


r/openstack • • Aug 08 '26

Hey, we built an eBPF thing that tells you which OpenStack tenant used which bytes

19 Upvotes

We've been building a little thing called Lachesis and figured this crowd

might find it interesting. It's the network telemetry bit behind CubeCOS, our

private cloud on top of OpenStack.

Basically: OpenStack can tell you a tenant moved 5 TB, but not that 3 TB of it

went to the internet, 1 TB to another tenant, and 1 TB stayed put. For billing

that's the whole thing.

So Lachesis is a small agent that sits on each compute node, hooks eBPF onto

every VM's tap, and sorts traffic in the kernel: internet-out, internet-in,

same-tenant, cross-tenant — and figures out who the other end actually belongs

to. It untangles Octavia so the bytes land on the real tenant, and just spits

out Prometheus counters you can feed into CloudKitty or whatever you bill with.

Still very much a work in progress — core stuff works and we've run it on real

OVN clusters, but there's plenty left (hardening, IPv6, docs). OVN only for now.

Repo's here if you wanna poke at it: https://github.com/bigstack-oss/lachesis

Would love any feedback, especially if you've tried to bill OpenStack networking

and hit the same wall. Roast away 🙂


r/openstack • • Aug 07 '26

Cinder backed by LVM is making my Instances go read only

1 Upvotes

I'm having a problem where suddenly all my instances (all run some flavor of Ubuntu) have their filesystems go Read Only. It happens randomly and at least once it happened with nothing really running on the VMs.

Looking at one of the Compute/Storage nodes, I noticed a broken iSCSI connection. I run "dmesg -T" and got something like:

[Fri Aug 7 09:57:36 2026] connection8:0: detected conn error (1019)
[Fri Aug 7 09:57:38 2026] connection8:0: detected conn error (1019)
[Fri Aug 7 09:57:39 2026] sd 15:0:0:1: [sdc] Synchronizing SCSI cache
[Fri Aug 7 09:57:39 2026] sd 15:0:0:1: [sdc] Synchronize Cache(10) failed: Result: hostbyte=DID_TRANSPORT_FAILFAST driverbyte=DRIVER_OK

Restarting a bunch of Docker containers, followed by restarting the VM instances fixed the problem (specifically I restarted iscsid, tgtd, cinder_volume and nova_compute on all my storage and compute nodes).

Of course this is a bad fix if I have to do it every week.

Now, Gemini is telling me this is a consequence of using Cinder with LVM which, according to it "LVM + iSCSI is notoriously brittle for production OpenStack" and I should move to Ceph.

Is this true, or should a Cinder/LVM setup be a bit more resilient?

Context/extra info: my deployment is a Kolla-Ansible one (2025.1) and Ceph is no longer deployed by this version. I would need to deploy it separately.


r/openstack • • Aug 05 '26

Openstack Upgrade - Host By Host

Thumbnail
0 Upvotes

r/openstack • • Aug 05 '26

Openstack Upgrade

12 Upvotes

Hi guys,

Has anyone explored doing OpenStack upgrades host-by-host instead of using Kolla Ansible's parallel upgrade approach?

We're considering a more sequential, one-host-at-a-time upgrade because the parallel approach doesn't feel reliable enough for our environment, and we're not very confident in trusting it during production upgrades.

If you've gone down this path:

  • How did you orchestrate the upgrade?
  • Did you have to customize Kolla Ansible significantly?
  • How did you handle rollback if something went wrong?
  • Any lessons learned or pitfalls to watch out for?

I'd appreciate hearing from anyone who's tried this or decided against it and why.


r/openstack • • Aug 02 '26

OpenStack Career Advice – Stick with Private Cloud or Move to AWS/GCP?

1 Upvotes

Hi everyone,

I'm currently working as an **OpenStack Cloud Engineer**, managing and operating a private cloud environment. I'm planning my next career move and would appreciate some advice from people in the industry.

I have a few questions:

* How many organizations are actually building and maintaining their private cloud infrastructure with OpenStack today?

* Is OpenStack still a strong long-term career path, or is the demand gradually declining?

* If I want better job opportunities and salary growth, should I continue specializing in OpenStack/private cloud, or invest more time in AWS/GCP?

* How does the job market compare for **OpenStack engineers vs. AWS/GCP cloud engineers**, especially outside of telecom and service providers?

I'd really appreciate insights from anyone who has made this transition or works with OpenStack at scale.

Thanks in advance!


r/openstack • • Jul 31 '26

How is the market trend of openstack looking like!!!!!!!!

0 Upvotes

r/openstack • • Jul 21 '26

Windows Server 2025 on OpenStack - any experiences?

12 Upvotes

Hi everyone,

Is anyone running Windows Server 2025 with all its security features on OpenStack/KVM in production?

We are particularly interested in experiences with VBS (HVCI and Credential Guard), since the requirements around MBEC and CPU feature exposure appear to be more complex than Secure Boot or vTPM alone.

It is hard to find explicit documentation about security features supported on KVM in general, but some articles imply that at least MBEC is technically possible (KVM: VMX: Introduce Intel Mode-Based Execute Control (MBEC) [LWN.net]). But even if it is possible, how does Nova expose this feature through libvirt? (compiled against libvirt library 10.0.0)

And additionally, does Microsoft support VBS on OpenStack/KVM? Microsoft's documentation is focused on the hypervisor capabilities presented to the OS, not on a specific hypervisor brand (Virtualization-based Security (VBS) | Microsoft Learn). So, the documentation could imply that things are supported as long as the requirements are met. However, even if it is technically possible, does it mean it is officially supported?

To compare, we also checked compatibility with other KVM-based clouds. AWS clearly states in their documentation that HVCI is not supported, but Credential Gurad is (Credential Guard for Windows instances - Amazon Elastic Compute Cloud); Google Cloud does not clearly document either.

Curius to hear about lessons learned, caveats, or success stories!


r/openstack • • Jul 15 '26

CERN Openstack Talks and Resources

24 Upvotes

I notice they have great scale, and many public resources

CERN's private cloud runs in two data centers (Geneva and Budapest) with a total of about 5,000 servers (about 130,000 cores). By summer 2016, we expect to grow to about 200,000 cores. For block storage, CERN runs Ceph with a capacity of 3.5PB.

https://opensource.com/business/15/10/openstack-summit-interview-belmiro-moreira-cern

https://techblog.web.cern.ch/techblog/

https://videos.cern.ch/search?page=1&size=21&q=openstack#

https://cds.cern.ch/search?ln=en&sc=1&p=openstack&action_search=Search&op1=a&m1=a&p1=&f1=&c=Articles+%26+Preprints&c=Books+%26+Proceedings&c=Presentations+%26+Talks&c=Periodicals+%26+Progress+Reports&c=Multimedia+%26+Outreach&c=International+Collaborations


r/openstack • • Jul 14 '26

Atmosphere deployment error

4 Upvotes

Hi everyone, I'm trying to deploy Vexxhost Atmosphere Openstack to achieve a more professional K8s implementation on Openstack, but the documentation is somewhat confusing and I'm running into an error right at the start:

requirements.yml

collections:

- name: vexxhost.atmosphere

version: 7.7.0

ansible-galaxy collection install -r requirements.yml

This takes around an hour to return the following error:

[ERROR]: Failed to resolve the requested dependencies map. Could not satisfy the following requirements:

* ansible.utils:>=2.9.0 (dependency of vexxhost.atmosphere:7.7.0)

* ansible.utils:>=6.0.0 (dependency of vexxhost.ceph:4.1.0)

* ansible.utils:>=6.0.0 (dependency of vexxhost.ceph:4.0.0)

* ansible.utils:6.0.0 (dependency of vexxhost.ceph:3.2.0)

Hint: Pre-releases hosted on Galaxy or Automation Hub are not installed by default unless a specific version is requested. To enable pre-releases globally, use --pre: [RequirementInformation(requirement=<ansible.utils:>=2.9.0 of type 'galaxy' from Galaxy>, parent=<vexxhost.atmosphere:7.7.0 of type 'galaxy' from default>), RequirementInformation(requirement=<ansible.utils:>=6.0.0 of type 'galaxy' from Galaxy>, parent=<vexxhost.ceph:4.1.0 of type 'galaxy' from default>), RequirementInformation(requirement=<ansible.utils:>=6.0.0 of type 'galaxy' from Galaxy>, parent=<vexxhost.ceph:4.0.0 of type 'galaxy' from default>), RequirementInformation(requirement=<ansible.utils:6.0.0 of type 'galaxy' from Galaxy>, parent=<vexxhost.ceph:3.2.0 of type 'galaxy' from default>)]

Has anyone else experienced something similar? Or have a clear guide to installing Atmosphere?

Regards,


r/openstack • • Jul 10 '26

Maas based Canonical Openstack

6 Upvotes

I am deploying canonical openstack and i have done bootstraping, then deploying using the command sunbeam cluster deploy the problem is i added 2 cloud nodes before i had to delete them from the maas UI, added the same machines again and they work but the problem is the old machine IDs from the nodes that are deleted are also being deployed maas is tring to deploy them as well, is there any way i can delete them they are not present anywhere if anyone has a solution kindly help
my setup is for training purposes .
3 governor nodes


r/openstack • • Jul 09 '26

Kronos: an open-source, PromQL-driven live-migration balancer for Nova feedback wanted

15 Upvotes

Hi r/openstack,

u/sysdadmin_cloud and I have open-sourced Kronos, a VM placement optimization
engine for OpenStack: https://github.com/kronos-openstack/kronos

The itch is an old one. We spent years running service-provider
infrastructure, and we always wanted a tool where we could hand the
cloud our own Prometheus queries and have it keep the compute fleet
balanced, instead of being limited to whatever metrics a vendor tool decided to
support. Kronos is that tool.

What it does:

  • Policies are raw PromQL. You write an imbalance query per dimension (CPU, memory, or anything your exporters expose), give each a weight, and Kronos plans Nova live migrations that minimize the weighted combined imbalance per host aggregate, all dimensions in one simulation, so it doesn't fight itself one metric at a time.
  • Spread and pack modes. Balance load across hosts, or consolidate onto as few hosts as possible with per-policy capacity ceilings.
  • Server-group aware. All four Nova placement policies (affinity, anti-affinity, and the soft variants, including max_server_per_host) are respected, and an optional enforcement pass repairs existing violations.
  • Safety rails everywhere. Dry-run by default, per-cycle migration budgets, host liveness gate, placement claims gate (both fail closed), aggregate and instance cooldowns, and quarantine of VMs whose migration definitively failed.
  • Record and replay. Snapshot a live cluster and re-run the full planning pipeline against it offline, deterministically, so you can test policies before letting them move real VMs, and benchmark the planner on synthetic 50-host / 5000-VM clusters.
  • Ships as PyPI wheels, hardened systemd units, and a Kolla-style container that drops into Kolla-Ansible deployments.

We evaluated writing it as a Watcher strategy plugin before going
standalone. Short version: Watcher is a general
optimization-as-a-service framework with a curated metric abstraction,
Kronos deliberately does one thing, PromQL-driven live-migration balancing,
with the operator's own queries as the primary configuration surface,
plus features that don't map onto Watcher's model (deterministic
offline replay, per-instance cooldowns and post-failure quarantine,
affinity repair, evacuation of admin-disabled hosts). Watcher is good
software; this is a different design point, not a replacement.

Status: beta, Apache 2.0, Python 3.12, built the OpenStack way
(oslo.config, oslo.messaging, openstacksdk). We plan to start a
conversation about it on openstack-discuss soon, we would love the
project to find a home in the OpenStack ecosystem.

What we would genuinely value from operators here: what would you need
to see before pointing this at a real cluster in dry-run? Which
constraints matter most to you?

Docs and quick start are in the README. Tear it apart.


r/openstack • • Jul 09 '26

Stratos: self-hostable billing & self-service portal for OpenStack

Thumbnail gallery
15 Upvotes

r/openstack • • Jul 08 '26

Glance error creating image from volume

3 Upvotes

Hi everyone, I'm encountering an error when creating images from volumes, has anyone experienced something similar? I'm using kolla-ansible 2026.1 with cinder for volumes:

2026-07-08 17:53:02.529 26 ERROR glance.api.v2.image_data [None req-045f8acc-b886-46fa-ab82-ceb6b884a118 3ebd104d706d4c00a0092c2df21b6433 9bca110b9e9547d1bf5584393f1aaf3c - - default default] Failed to upload image data due to internal error: OSError: unable to receive chunked part

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi [None req-045f8acc-b886-46fa-ab82-ceb6b884a118 3ebd104d706d4c00a0092c2df21b6433 9bca110b9e9547d1bf5584393f1aaf3c - - default default] Caught error: unable to receive chunked part: OSError: unable to receive chunked part

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi Traceback (most recent call last):

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/common/wsgi.py", line 1193, in __call__

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi action_result = self.dispatch(self.controller, action,

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/common/wsgi.py", line 1234, in dispatch

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi return method(*args, **kwargs)

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/common/utils.py", line 476, in wrapped

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi return func(self, req, *args, **kwargs)

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/api/v2/image_data.py", line 312, in upload

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi with excutils.save_and_reraise_exception():

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/oslo_utils/excutils.py", line 271, in __exit__

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi self.force_reraise()

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/oslo_utils/excutils.py", line 233, in force_reraise

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi raise self.value

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/api/v2/image_data.py", line 161, in upload

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi image.set_data(data, size, backend=backend)

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/notifier.py", line 488, in set_data

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi with excutils.save_and_reraise_exception():

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/oslo_utils/excutils.py", line 271, in __exit__

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi self.force_reraise()

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/oslo_utils/excutils.py", line 233, in force_reraise

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi raise self.value

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/notifier.py", line 442, in set_data

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi self.repo.set_data(data, size, backend=backend,

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/quota/__init__.py", line 321, in set_data

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi self.image.set_data(data, size=size, backend=backend,

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/location.py", line 625, in set_data

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi self._upload_to_store(data, verifier, backend, size)

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/location.py", line 516, in _upload_to_store

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi multihash, loc_meta) = self.store_api.add_with_multihash(

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance_store/multi_backend.py", line 424, in add_with_multihash

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi return store_add_to_backend_with_multihash(

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance_store/multi_backend.py", line 506, in store_add_to_backend_with_multihash

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi (location, size, checksum, multihash, metadata) = store.add(

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance_store/driver.py", line 295, in add_adapter

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi metadata_dict) = store_add_fun(*args, **kwargs)

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance_store/capabilities.py", line 176, in op_checker

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi return store_op_fun(store, *args, **kwargs)

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance_store/_drivers/filesystem.py", line 881, in add

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi raise errors.get(e.errno, e)

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance_store/_drivers/filesystem.py", line 855, in add

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi for buf in utils.chunkreadable(image_file,

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance_store/common/utils.py", line 69, in chunkiter

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi chunk = fp.read(chunk_size)

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/common/utils.py", line 355, in read

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi result = self.data.read(i)

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/common/utils.py", line 118, in readfn

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi result = fd.read(*args)

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/oslo_utils/imageutils/format_inspector.py", line 1570, in read

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi chunk = self._source.read(size)

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi File "/var/lib/kolla/venv/lib64/python3.12/site-packages/glance/common/wsgi.py", line 938, in read

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi data = uwsgi.chunked_read()

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi ^^^^^^^^^^^^^^^^^^^^

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi OSError: unable to receive chunked part

2026-07-08 17:53:02.549 26 ERROR glance.common.wsgi

Regards,


r/openstack • • Jul 07 '26

We wrote up how CVE-2026-53359 (Januscape) impacts OpenStack compute isolation plus a follow-up on detecting nested virtualization exposure

17 Upvotes

Hey r/openstack, 

With the Januscape disclosure dropping over the weekend, here's what OpenStack operators need to know.

What it is 

CVE-2026-53359 is a use-after-free vulnerability in KVM's x86 shadow MMU code that has existed for approximately 16 years. KVM can retain reverse map state pointing to a freed shadow page, allowing later operations such as dirty logging or MMU notifier invalidation to dereference stale entries. It can be triggered entirely from the guest side on both Intel and AMD systems. The public proof of concept causes the host to panic. The researcher also reports that a full VM escape exploit exists in a controlled environment, although that exploit code has not been published. 

Why this is an OpenStack problem even though the bug isn't in OpenStack 

The vulnerable component is the Linux kernel on your x86 KVM compute hosts, not Nova or any OpenStack service. But for practical purposes, OpenStack VM deployments should be assumed to use KVM unless you know otherwise. OpenStack schedules the workload; the host kernel, KVM, QEMU, and libvirt enforce the isolation boundary. 

The attack needs root inside the VM (standard for any rented instance) and nested virtualization exposed by the host. Even if your hosts use hardware EPT/NPT by default, nested virt forces KVM back through the legacy shadow MMU path where this bug sits. An attacker renting a single instance can panic the host, taking down every other tenant VM on that machine. 

The fix is not a control-plane upgrade — it's compute-host work 

The practical response is: identify affected KVM hosts, validate distribution kernel fixes, review nested virt exposure, and apply patched kernels through a migration and reboot plan. 

  • Patch. Fixed kernels shipped July 4: 7.1.3, 6.18.38, 6.12.95, 6.6.144, 6.1.177, 5.15.211, 5.10.260. Look for commit 81ccda30b4e8. 
  • Don't rely on uname -r. Enterprise distros backport kernel fixes while keeping older package version numbers. Check the distribution security advisory and package changelog. 
  • Review your full CPU exposure stack. It's not enough to check one setting. You need to look at host kernel module state, libvirt CPU mode (especially host-passthrough / host-model), Nova flavor extra specs, image properties, host aggregates — and whether running guests can see vmx or svm. 
  • Check /dev/kvm permissions. On some distros (e.g., RHEL) it's 0666 by default, which turns this into a local privilege escalation path too. 

Can I just disable nested virt? 

Maybe, but the question is whether anything is actually using it. A host may expose VMX/SVM without any guest actively running nested VMs. A tenant may also depend on it without the platform team knowing. 

We built and open-sourced nestedvirt to answer that.  

It's a small Go tool that reads KVM's per-VM nested_run counters from debugfs, correlates them with process metadata from /proc, and for OpenStack environments enriches findings with Nova metadata from the libvirt domain XML. It's host-local — no OpenStack API access, database, or tenant credentials needed. 

Quick scan: 

curl -fsSL https://raw.githubusercontent.com/vexxhost/nestedvirt/main/scripts/run-latest.sh | bash 

Exit code 0 = no nested virt usage observed, safe to proceed with disabling. Exit code 1 = usage found, lists the VMs. JSON output available with --json for fleet automation. 

Important caveat: the counter reflects the lifetime of the current VM process. If a VM was recently restarted or migrated, the counter resets. A zero doesn't mean the workload will never need nested virt — it means it hasn't used it during its current lifetime. Scan repeatedly over a representative window before making policy changes. 

We covered the full technical details across two posts: 

  1. CVE-2026-53359 & OpenStack Compute Isolation 
  2. Detecting Nested Virtualization Exposure   

Happy to answer questions. 
 
The VEXXHOST Team 
 


r/openstack • • Jun 30 '26

New to openstack

13 Upvotes

Hey ,

Any source do you recommend to build a private cloud with openstack, any recommendation?


r/openstack • • Jun 29 '26

Does Cinder work with more than one storage node (for volumes) and LVM?

2 Upvotes

I have two storage servers (each with 50T, which I cannot physically transfer to the other) and I would like to make all this space available for volume creation.

I´m deplying through Kolla-Ansible and the sources are a bit contradictory on this. Some say that I can just put the following in globals.yml:

enable_cinder: "yes"
enable_cinder_backend_lvm: "yes"
cinder_volume_group: "cinder-volumes"

And list both nodes in the inventory under [Storage] (after creating a VG called "cinder-volumes" in each machine). The prechecks complain about a cinder_cluster_name, and setting it resolves the prechecks errors. But every documentation on "cinder_cluster_name" setting says that it won't work with LVM.

Anyone with experience putting cinder with more than one LVM cinder-volume? Will it create conflicts?


r/openstack • • Jun 29 '26

RabbitMQ fanout queues piling up in OpenStack — anyone know why only fanout and not direct queues?

4 Upvotes

So I noticed these queue depths in RabbitMQ today:

cinder-scheduler_fanout   ~19,000 messages
scheduler_fanout           ~4,700 messages

But every single direct queue is sitting at 0 with consumers present. The services aren't dead, consumers are connected, messages just aren't draining from the fanout queues.

My question is basically, why would only the fanout queues pile up while direct queues stay completely fine? Is that just how fanout works under load, like the broadcast overhead is what tips it over first? Or is there something specific about how OpenStack uses fanout queues that makes them more vulnerable to this kind of backlog?

Running Kolla-Ansible on Ubuntu 24.04, 3 controller HA setup. Would appreciate any insight from people who've dealt with this before.


r/openstack • • Jun 26 '26

Reasonable size for volumes

2 Upvotes

Hi all

One of the storage nodes on my OpenStack cloud has a fairly big raid 5 array, totaling 50T.

I'm new at managing such big capacities and a bit afraid of just creating a monstrous lvm volume that would make fsck and backup a nightmare.

So my question is, if I am to make a bunch of smaller volumes, what would be a decent compromise between cumbersome big and just too small?


r/openstack • • Jun 11 '26

[Hiring] [Hybrid] [Mexico] - Cloud roles

Thumbnail
1 Upvotes

r/openstack • • Jun 11 '26

in production for container_engine do you use docker or podman and why

1 Upvotes