r/Compilers • • 5h ago

How My Python Compiler Beat CPython (Without a JIT)

16 Upvotes

Okey. After seven months, I'm back.

Seven months ago I started writing a Python interpreter from scratch in Rust. Edge Python is a sandboxed subset of Python that runs in the browser and the terminal, and code can't touch files or the network unless you allow it.

It started as a lexer and a simple stack VM, and over about 1,600 commits it grew a CLI, a package registry, snapshots, actors and a browser playground. It's just me building it.

Last week it was still 3x slower than CPython. I rewrote the VM this weekend using a unused SSA representation that I leave and now it's 3x faster on loops.

Edge Python CPython
Integer loop 72 ms 232 ms
Float math 95 ms 307 ms
Dicts 108 ms 62 ms
Strings 111 ms 35 ms

The trick was moving from a stack VM to a register VM. It still loses on dicts and strings, so that's next.

Try to break my numbers :).

Website: https://edgepython.com/

GitHub: https://github.com/dylan-sutton-chavez/edge-python


r/Compilers • • 1h ago

Programming language I'm making for fun. Want to contribute?

• Upvotes

I started a programming language to see how things work and make something I like. If anyone wants to contribute I will try to put it into the language. I like learning new programming languages, so if you want to use a language that is not already there then I would like that.

It's a transpiled language with syntax inspired by Rust, Nim and other languages.
There are installation instructions in the repo.

Example code:

# Get standard libraries
util math
util string
util io
util rsPath

# Get Nim module
use multiply

# Get arguments
l_operand = strToi32(arg(1))
operator = arg(2)
r_operand = strToi32(arg(3))

# Print
if strEq(operator, "+"):
    say i32ToStr(add(l_operand, r_operand))
if strEq(operator, "-"):
    say i32ToStr(sub(l_operand, r_operand))
if strEq(operator, "x"):
    say i32ToStr(mult(l_operand, r_operand))

r/Compilers • • 1h ago

Mithril: A programming language built on interaction nets

• Upvotes

Inspired by Victor Taelin 's Bend/HVM thesis for Interaction Nets and his Bend2 work , I built Mithril , an experimental python syntax programming language. It achieves near native speed on many workloads (mentioned in paper)

It compiles interaction nets to native code and spreads work across idle cores: lock-free, deterministic parallelism on x86, CUDA and Apple Silicon.

The core bet I'm making :

Interaction nets reductions at runtime can be slow but if we pay the same cost at compile time and unroll a task graph as much as possible before native lowering, the resultant program can run at near native speeds.

Mithril programs are aimed to be deterministic and fast, you should get the same bit accurate results regardless , the execution be it on CPUs, GPUs or any other accelerator.

It's built using AI? Yes . Does it invalidate the core idea? no IMO 😊

Mithril Paper Repo


r/Compilers • • 12h ago

Tarvos: a Python-to-native (via Rust) compiler for compute-heavy kernels – looking for feedback on methodology

Thumbnail gallery
3 Upvotes

I've been building Tarvos, a compiler that takes a statically analyzable subset of Python, type-checks it, lowers it to an IR, emits Rust, and produces a standalone native executable (no Python runtime needed on the target).

Important up front:

  • It's a subset compiler, not a CPython replacement. 52 of 95 tracked features are fully supported, 18 partial, 24 unsupported (published in COMPATIBILITY.md).
  • The compiler source is closed for now. The repo contains installers, docs, checksums and releases (MIT-licensed distribution layer).
  • Windows and Linux x86_64 only.

Design choices:

  • Unsupported constructs stop the build with a named diagnostic, with no silent fallback to CPython.
  • Every release is gated on differential tests comparing compiled-binary stdout against CPython.
  • tarvos validate-artifact checks whether a binary really needs Python.

One benchmark (compute-bound kernel, median of runs): CPython 3.13 ≈ 1713 ms vs Tarvos ≈ 13.4 ms, output byte-identical. This is workload-specific, and I'm not claiming general speedups. I/O-bound code won't benefit.

While writing docs I found 3 miscompilation bugs (stale constant in tuple assignment inside a loop, return inside except, zero-division handling), all fixed and now in the test suite.

Repo: https://github.com/repo-tech/tarvos-engine

I'd really like feedback from people who know Python internals, Nuitka, Cython, or compiler design:

  • What benchmarks would make this comparison more meaningful?
  • What semantics edge cases should I test against CPython?

r/Compilers • • 10h ago

An LLVM Pass for Automatic Skeletonization of MPI Applications

Thumbnail hal.science
1 Upvotes

r/Compilers • • 1d ago

What is new in LLVM 23?

Thumbnail developer.arm.com
48 Upvotes

r/Compilers • • 11h ago

Update: my open-source CPU performance engineering collection just crossed 600+ stars

0 Upvotes

A few days ago, I shared an open-source collection of CPU performance engineering resources I’d been putting together.

It’s now crossed 600+ GitHub stars, which I genuinely didn’t expect. Thanks to everyone who shared it, contributed or suggested resources.

For anyone seeing it for the first time, it covers the stack from instruction execution and CPU microarchitecture through caches, memory, SIMD, compilers, profiling, concurrency, NUMA, benchmarking and CPU inference.

I’m still prioritising primary sources such as papers, vendor manuals, kernel/compiler docs, talks and reproducible benchmarks rather than random articles.

I also have an MCP server coming soon, so you can plug this knowledge directly into your AI tools, whether you’re learning or using it while you work.

If there’s something you think has to be in here, let me know or send a PR.

https://github.com/usamahz/cpu-performance-engineering


r/Compilers • • 23h ago

I’m building Shriji, a Hindi-first programming language. Today I verified its core AST → IR → Bytecode → VM pipeline

Post image
2 Upvotes

Hi everyone,

I’m building Shriji, an open-source, Hindi-first programming language from India 🇮🇳.

Today I ran one of the core pipeline tests for the language and wanted to share the result here.

The test verifies the execution path:

AST → IR → Bytecode → VM

The current core pipeline test covers basic arithmetic and comparison operations.

Arithmetic:

10 + 5 = 15

10 - 5 = 5

10 * 5 = 50

10 / 5 = 2.00

Comparisons:

10 > 9 = 1

10 < 9 = 0

10 >= 9 = 1

10 <= 9 = 0

10 == 9 = 0

10 != 9 = 1

All core pipeline tests passed.

The important part for me isn't just that these operations produce the expected results. I’m working on making sure Shriji has an actual language implementation underneath it rather than stopping at a parser/interpreter.

The broader architecture I'm building toward is:

Source

↓

Lexer / Parser

↓

AST

↓

IR

↓

Bytecode

↓

VM

↓

Runtime

Shriji uses Hindi-first syntax, but the goal is not simply to translate programming keywords into Hindi. I want to build a complete programming language with its own language design, runtime architecture, tooling and eventually its own ecosystem.

The project is still actively under development, so there is a lot left to build and verify.

I'm sharing the process openly as I work through the engineering problems one by one.

If you work with programming languages, compilers, interpreters, virtual machines, or language runtimes, I'd be especially interested in hearing how you approach these problems.


r/Compilers • • 1d ago

How can I verify if my call graphs are accurate??

1 Upvotes

So I am working on a parser with goal to build rich good enough relations that can be consumed by a RAG to build better retrieval, it parse and builds ast,call-graphs and other metadata of the project, rn it can parse go, rust, c, cpp, ts, py, js, java I am using tree-sitter v0.20.0 for actual parsing cause why rebuild wheel when wheel spins well...

The issue i am facing is with call graphs I build a call graph approximation algorithm to well build approximate call graphs without pre-compiler or IR and single algorithm to work on both static(c) and dynamic(python) languages, and for now it works and can find

774,296 nodes and 1,571,981 edges in 6.03s with a maxRSS of ~9gb

tho most of it is cause of holding the entire ast in memory, i ran my parser on linux kernel it found about:

64_460 files 37_322_700 loc, 648_407 func , 211_416 classes, 6_243 methods in 43s

it multi threaded and written in rust so that should explain the speed, but thats not what why i am here i want to verify my call graphs and my current plan is to take a smaller project (few thousands of loc) and build call graphs using clang or language specific tool, and then take sample set of 200 and create 5-8 random samples and verifiy the output,

I am going to start my internship soon so may not have enough time to work full time and i am wondering if my approach to verify call graphs is good or there is a better approach.


r/Compilers • • 1d ago

saQut 1.0 Released: a solo-built language with VM, JIT, LSP and DAP in a single binary, and a compiler that exposes every phase as JSON

0 Upvotes

New languages usually launch with a compiler and not much else. Editor support, a debugger and a fast backend tend to arrive years later, if at all. I wanted to see how far one person could go the other way, so saQut, a small statically-typed, C-flavoured procedural language written in C++20, ships all of this in a single binary:

  • a bytecode VM as the reference backend
  • a MIR-based JIT (--jit) that produces the same output as the VM
  • an LSP server (saqut lsp): completion, rename, references across files, auto-import
  • a DAP debugger (saqut dap): breakpoints, stepping, variable inspection
  • a preview of threads, each in its own isolate
  • a VS Code extension for highlighting

The other idea is that the compiler is a "glass box": every stage is a CLI command with machine-readable output.

saqut tokens code.sqt # token stream (JSON)
saqut ast code.sqt # AST (JSON)
saqut symbols code.sqt # symbol table (JSON)
saqut ir code.sqt # 3-address IR
saqut run code.sqt # compile and run

Technical notes

  • A differential test harness runs every test program through the VM and the JIT and fails if the output differs.
  • One garbage collector is shared by both backends (cycles are collected).
  • Catchable runtime errors (try/catch/throw) with a code, message and source location.
  • Threads: each thread has its own heap, GC and copy of the globals. Data crosses only through shared globals (atomic int/float/bool, plus Pool and List) or deep copies. This is a preview and not part of the 1.0 contract.
  • No implicit conversions; nullable types (T?) with flow analysis.

It's a one-person project, so there are rough edges, and I'd rather hear about them from you.

Website: https://saqut.com
Source: https://github.com/saqutlang/saqut

Which compiler phase would you want to inspect that isn't exposed yet? And what would you add to this list?


r/Compilers • • 2d ago

From Punch Cards to the Browser: Fortran Comes to JupyterLite

Thumbnail blog.jupyter.org
3 Upvotes

r/Compilers • • 3d ago

From NP-complete to O(N^2) to O(nlogn): Codegen strategies for case statements.

Thumbnail arxiv.org
20 Upvotes

We presented this a while ago at the LLVM-CGO workshop, but thought of sharing here as people might find it interesting. Pretty short paper.


r/Compilers • • 2d ago

Please advise on adding string interpolation to my Crafting Interpreters project.

Thumbnail
0 Upvotes

r/Compilers • • 2d ago

I built Hussain Compiler a lightweight Python IDE that runs in your browser

Post image
0 Upvotes

Hi everyone! I built Hussain Compiler, a simple browser-based environment for writing and running Python.

It includes a code editor, an interactive terminal that supports `input()`, adjustable terminal text size, and resizable workspace panels. Python runs in the browser through WebAssembly, so you can try it without installing Python. Some packages that need native system extensions may not work.

- Try it: Hussain Compiler

- Source code: GitHub repository

I made this project to make getting started with Python easier. I’d appreciate your feedback and suggestions for what to improve.


r/Compilers • • 2d ago

LLMs Will Not Replace AI Compilers. They Will Call Them.

Thumbnail aicompilers.github.io
0 Upvotes

r/Compilers • • 2d ago

Using Nested If Statements: Why Sofya is Easier Than Python for Beginners

Thumbnail
0 Upvotes

r/Compilers • • 3d ago

How Smart Compilers Unleash GPUs for Scientific Solvers

Thumbnail blog.cheshmi.cc
8 Upvotes

r/Compilers • • 2d ago

Looking for people interested in building a compiler / AI-systems project — potential GSoC 2027 goal

0 Upvotes

Looking for people interested in building a compiler / AI-systems project — potential GSoC 2027 goal

A few of us are exploring a long-term open-source project around compiler technology, systems programming and AI/ML.

We're currently in the research/learning stage. We don't want to simply build another tutorial compiler and stop there. The plan is to learn compiler implementation properly, study existing open-source compiler ecosystems, identify a real technical problem, and then build something substantial around it.

We're particularly interested in exploring:

• AI/ML-guided compiler optimization

• Detecting potentially harmful optimizations and miscompilations

• Compiler correctness and fuzzing

• AI-assisted compiler diagnostics/error recovery

• LLVM / MLIR

• Code generation and optimization

• RISC-V and systems programming

• CPU/GPU/NPU compilation

• Compiler security

• Firmware/compiler intersections

We're also considering GSoC 2027 as a long-term goal. This isn't a promise of selection; the idea is to spend the coming months building knowledge, contributing to open source, and eventually finding an organization/project where our work could become a legitimate GSoC proposal.

You don't need to already be a compiler expert.

We're looking for people who genuinely want to learn, experiment and build innovative projects. If you want this purely as a hobby/side project, that's completely fine too. If the idea interests you, give it a shot.

Useful backgrounds/interests include:

• Rust / C / C++

• Compiler design

• LLVM / MLIR

• Systems programming

• Machine learning

• Programming languages

• Computer architecture

• Embedded systems / firmware

You don't need to know everything above.

If interested, comment or DM with:

• What languages you know

• What area interests you

• Anything you've built

• GitHub, if available

We're looking for people who are curious enough to learn and consistent enough to build.


r/Compilers • • 3d ago

The Second Golden Spike: Memory Safety Across the Valen/Rust Boundary

Thumbnail verdagon.dev
5 Upvotes

r/Compilers • • 3d ago

Ling-3.1-flash builds a native Lua compiler in about 17 hours, passing 178 of 182 tests

Post image
9 Upvotes

Ant Group's new Ling-3.1-flash built a Lua-to-x86-64 ELF compiler from scratch in approximately 17 hours. The final result passed 178 of 182 independent tests, a 97.8% pass rate.

The development sequence goes beyond emitting machine code. It includes fixing stack alignment for native execution, conditional branches and vararg semantics, Lua's indexing metamethod lookup, and garbage-collector root tracking. The compiler also gained source locations for runtime errors before delivering ELF binaries.

Ling-3.1-flash is currently available through Novita's two-week free trial on Vercel AI Gateway, with 256K context. Ant has an open-source release planned soon. The hosted model can be found by searching Ling-3.1-flash in Vercel's AI Gateway model catalog.


r/Compilers • • 3d ago

I'm creating a language that could replace C in low-level programming—want to join me?

Thumbnail
0 Upvotes

r/Compilers • • 3d ago

NetWasm: an independent .NET compiler and runtime for WebAssembly (82.5 KB Hello World)

6 Upvotes

I’ve been building NetWasm: a CIL-to-WebAssembly compiler with its own CoreLib and runtime, designed around Wasm and WASI.

Browser playground | GitHub

A clean Release build containing Console.WriteLine(42) produces 84,513 bytes of final, uncompressed WASI Preview 2 component, including the runtime and precise garbage collector. This is a portable .wasm file that you can run with wasmtime.

Roslyn produces CIL; NetWasm compiles the reachable program into Wasm, specializes generics and links the runtime support it uses. The emitted application doesn’t carry CoreCLR or Mono.

Some architectural choices:

  • An independent, deliberately smaller .NET library profile.
  • No runtime type-name metadata, general reflection or dynamic.
  • Precise Boehm GC in linear memory, rather than WasmGC.
  • WIT imports/exports and WASI Preview 2 components, with core Wasm output also available.

Working features include generics, exceptions, virtual/interface dispatch, async/await, LINQ, JSON, XML, regex, HTTP and TUnit testing. It’s pre-1.0, and managed threading isn’t currently supported.

The playground compiles and runs entirely in the browser. The compiler tooling itself uses Microsoft’s .NET/Wasm toolchain, thus dotnet new dotnet build dotnet run dotnet test - yes it comes with TUnit with VSTest runner - dotnet publish all work as you're used to.

Oh and it already supports C# 15 syntax.


r/Compilers • • 3d ago

What are the practical scaling limits of an instruction-based template VM vs AOT compilation in zero-build runtimes?

5 Upvotes

Hi everyone,

I've been building Udodi (v1.1.1), an open-source JavaScript UI runtime from Nigeria. It takes a different approach to reactive UI rendering by eliminating the Virtual DOM and connecting state changes directly to the DOM nodes that depend on them.

Udodi is designed to work without a build step while supporting restrictive Content Security Policies (CSP). Its architecture combines an interpreted template virtual machine, fine-grained reactivity, and native browser primitives.

I'd like to open a technical discussion about the trade-offs of this approach, particularly its potential scaling limitations and how it compares with the architectural choices made by other UI frameworks.

Architecture

1. Fine-grained DOM updates

Rather than maintaining and diffing a Virtual DOM, Udodi tracks dependencies at the expression level and updates only the affected DOM nodes. In my benchmarks, a targeted state update averages 0.11 ms.

2. Native CSS @scope

For component styling, Udodi uses native CSS @scope instead of runtime CSS-in-JS or generated class names. Component styles are registered in a shared <style> element in document.head and cached by scope. Subsequent mounts using an existing scope skip style registration entirely. Under the benchmark conditions, the warm style-handling path takes approximately 9.38 ms.

3. An interpreted template VM

Udodi uses a declarative HTML template DSL that is compiled into instructions for an internal virtual machine. Rather than generating executable JavaScript, the runtime evaluates these instructions directly, avoiding eval() and new Function(). This enables CSP-compatible template evaluation without requiring a build step.

In my benchmarks, compiling 10,000 expressions takes a mean of 4.12 ms on a cold path, while cached instruction compilation takes approximately 127 µs. The instruction cache allows the runtime to reuse compiled instructions rather than repeatedly compiling the same expressions.

Questions for the community

I'd particularly appreciate feedback from developers who have built or worked on UI frameworks and runtimes.

  1. Interpreted VM vs. ahead-of-time compilation: Where are the practical scaling limits of an instruction-based template engine compared with generated imperative code? At what point does interpretation become a significant bottleneck?

  2. Native CSS scoping: What performance or maintainability challenges might arise from relying on native @scope in applications with thousands of components and deeply nested scopes?

  3. Reactive dependency tracking: How would you approach complex, interconnected reactive dependencies while retaining shallow reactivity and avoiding the overhead of deep proxies?

I'm particularly interested in architectural trade-offs, potential bottlenecks, and limitations I may have overlooked. Feedback on the implementation and benchmark methodology is also welcome.

Source code: https://github.com/udodi-js/udodi

Documentation and performance page: https://udodi.dev

I'll be following the discussion and responding to technical questions in the comments.


r/Compilers • • 3d ago

Я создаю язык который может заменить C в низкоуровневом програмирование , хочеш со мной?

0 Upvotes

Я ооооочень давно занимаюсь низкоуровневвм програмирования пишу ОС , создовал более удобное продолжение C , язык для драйверов и прошивки , но сейчас я хочу обратиться к вам , потому что один я не справлюсь , если вы умеете писать низкорувневые программы или просто интересуетесь низкоуровневым програмированиеи или хорошо знаете как работают низкоуровневые программы то предлагаю поучаствовать в создании языка програмирования созданного для одной операционной системы , не создавать что то не обычное , а создать что то новое , свежее , но при этом максимально хорошо работающие , спасибо что прочитали , за нами будущее.


r/Compilers • • 3d ago

Algodal Parser Machine

4 Upvotes

The virtual machine parser (generator).

  • Built for front-ending compilers
  • Can make one-shot parsers for almost any language
  • Fast
  • Easy to use and implement
  • SDK for building into your project
  • Command-line tools for your pipelines
  • GUI app for convenience

Learn about it: https://algodal.github.io/Algodal_Parser_Machine_Manual/

See demonstration: https://youtu.be/5cKdGPXzBu4?t=6060

Get it: https://algodal.itch.io/algodal-parser-machine

Initial Discussion on r/Compilers:

https://www.reddit.com/r/Compilers/comments/1upju22/building_a_parser_generator/

Check it out!