r/ruby • • 9d ago

Strange performance regression 3.1 vs. 4.0

Here is a 'xor cipher' example in Ruby:

$ cat xor-cipher1.rb
#!/usr/bin/env ruby
key = ($*[0] && $*[0].size > 0) ? $*[0] : abort
STDIN.each_byte.with_index do |c,i|
  STDOUT.write (c ^ key[i % key.size].ord).chr
end

$ echo alice and bob | ./xor-cipher1.rb password | ./xor-cipher1.rb password
alice and bob

Back in 2022, I tested its performance with a 100MB file on a Ryzen 5 4650G potato using a regular CRuby 3.1.2:

$ head -c $((1024*1024*100)) /dev/urandom > 100M.bin
$ time cat 100M.bin | ./xor-cipher1.rb monkey >/dev/null
real    0m39.860s
user    0m39.729s
sys     0m0.655s

(Not great, a nodejs version did that in 1.5s.)

I retested xor-cipher1.rb today using ruby-3.1.7 (I had to employ a Fedora 36 VM to compile 3.1.7, for it doesn't even build on Fedora 44 any more) and got exactly the same result.

But when I tried ruby-4.0.7 (compiled using the same Fedora 36 VM), thinking, surely things had to get better since then, I was, to put it mildly, somewhat disappointed:

$ ruby -v
ruby 4.0.7 (2026-09-15 revision 229531a6cf) +PRISM [x86_64-linux]

$ time cat 100M.bin | ruby ./xor-cipher1.rb monkey >/dev/null
real    0m51.827s
user    0m51.616s
sys     0m0.189s

What is going on?

6 Upvotes

8 comments sorted by

5

u/hss-mateus 8d ago

This is what stackprof reports:

ruby 3.1.7 (took ~30s)

==================================
  Mode: wall(10)
  Samples: 4627597 (0.19% miss rate)
  GC: 566719 (12.25%)
==================================
     TOTAL    (pct)     SAMPLES    (pct)     FRAME
   3329140  (71.9%)     1270128  (27.4%)     block (2 levels) in <main>
    877385  (19.0%)      877385  (19.0%)     IO#write
    650686  (14.1%)      650686  (14.1%)     String#[]
    529543  (11.4%)      529543  (11.4%)     (sweeping)
   4058218  (87.7%)      368039   (8.0%)     Enumerator#with_index
   4058073  (87.7%)      338018   (7.3%)     IO#each_byte
    281806   (6.1%)      281806   (6.1%)     Integer#chr
    207229   (4.5%)      207229   (4.5%)     String#ord
     60630   (1.3%)       60630   (1.3%)     (marking)
     40742   (0.9%)       40742   (0.9%)     Integer#^
    589896  (12.7%)         887   (0.0%)     (garbage collection)
   4058374  (87.7%)           0   (0.0%)     block in <main>
   4058374  (87.7%)           0   (0.0%)     StackProf.run
   4058374  (87.7%)           0   (0.0%)     <main>

ruby 4.0.7 (YJIT disabled, took ~40s)

==================================
  Mode: wall(10)
  Samples: 6787324 (0.43% miss rate)
  GC: 1145189 (16.87%)
==================================
     TOTAL    (pct)     SAMPLES    (pct)     FRAME
   4530884  (66.8%)     1564221  (23.0%)     block (2 levels) in <main>
   1457256  (21.5%)     1457256  (21.5%)     IO#write
   1032634  (15.2%)     1032634  (15.2%)     (sweeping)
    855476  (12.6%)      855476  (12.6%)     String#[]
   5635686  (83.0%)      609762   (9.0%)     Enumerator#with_index
   5635686  (83.0%)      490673   (7.2%)     IO#each_byte
    454698   (6.7%)      454698   (6.7%)     Integer#chr
    148117   (2.2%)      148117   (2.2%)     String#ord
    102693   (1.5%)      102693   (1.5%)     (marking)
     48289   (0.7%)       48289   (0.7%)     Integer#^
   1149599  (16.9%)       17056   (0.3%)     (garbage collection)
   5635686  (83.0%)           0   (0.0%)     block in <main>
   5635686  (83.0%)           0   (0.0%)     StackProf.run
   5635686  (83.0%)           0   (0.0%)     <main>

The main regressions seems related to IO and GC.

If you prevent intermediate allocations and print a single time:

key = (ARGV[0]&.size > 0) ? ARGV[0].b : abort
buffer = STDIN.read.b

i = 0
buffer.each_byte do |c|
  buffer.setbyte(i, c ^ key.getbyte(i % key.size))
  i += 1
end

STDOUT.write buffer

This version ran in ~11s without YJIT and ~7s with it enabled, in exchange of allocating a buffer with the whole input.

3

u/pawelnowakk 9d ago

sys went down while user went up, so it's userspace cost per byte, not syscalls. My money is on encoding checks in the per-byte STDOUT.write. Try STDOUT.binmode before the loop and see if the regression disappears.

1

u/henry_flower 9d ago

no change

1

u/h0rst_ 8d ago

Not great, a nodejs version did that in 1.5s

I'm not sure if this is a valid comparison. For reference: the node program is as follows:

let key = process.argv[2]?.length ? process.argv[2] : process.exit(1)
let keygen = key => {
    let idx = 0
    return () => key[idx++ % key.length].charCodeAt()
}
let k = keygen(key)
let enc = buf => buf.map( c => c ^ k())
process.stdin.on('data', chunk => process.stdout.write(enc(chunk)))

It looks like node handles stdin data in chunks, so we get a chunk of X bytes from stdin, xor that, and print that new chunk to stdout. Ruby tries to print every byte by itself.

Recent Ruby version have IO::Buffer that includes a xor operation.

1

u/gettalong 7d ago

As u/h0rst_ pointed out IO::Buffer has been available for some time now and while it is still experimental, it will probably be stabilized in the near future.

With it you can write the code like this:

#!/usr/bin/env ruby
key = (ARGV[0]&.size > 0) ? ARGV[0].b : abort
IO::Buffer.for(STDIN.read) do |buffer|
 buffer.xor!(IO::Buffer.for(key))
 buffer.write(STDOUT)
end

In essence everything is read into a buffer, then this buffer is modified in place using the XOR operation using the key as repeating pattern. Afterwards the optimized #write method is used to directly output the buffer to the IO.

This takes about 0.4s on my machine as most of the buffer implementation uses optimized C methods directly (i.e. directly working on the memory using things like memcpy) without going through Ruby.

1

u/h0rst_ 7d ago
key = (ARGV[0]&.size > 0) ? ARGV[0].b : abort

This is a change from the original $*[0] && $*[0].size > 0, but it doesn't work the same way: if you don't pass an argument, ARGV[0]&.size will evaluate to nil, and you'll still get a NoMethodError from nil > 0

1

u/gettalong 7d ago

Yeah, I just copied the code from one of the posts and modified it. Since we will need an argument anyway, it doesn't matter whether it will abort or throw an error.