Slide 1

Slide 1 text

Hiroya Fujinami (a.k.a. makenowjust) @RubyKaigi 2026 in Hakodate / 2026-04-24 (Re)make Regexp in Ruby: Democratizing internals for the JIT

Slide 2

Slide 2 text

Hiroya Fujinami a.k.a. makenowjust • https://github.com/makenowjust / https://x.com/make_now_just • Ph.D. student at NII (National Institute of Informatics). • Studying the application of automata theory and formal languages. • Recently interested in nominal sets. • Currently on the job market (!) 2

Slide 3

Slide 3 text

Ruby committer

Slide 4

Slide 4 text

Regexp memoization

Slide 5

Slide 5 text

https://www.ruby-lang.org/ja/news/2022/12/25/ruby-3-2-0-released/

Slide 6

Slide 6 text

Regexp is good.

Slide 7

Slide 7 text

[YZ]JIT is good.

Slide 8

Slide 8 text

Regexp × JIT is good2...?

Slide 9

Slide 9 text

Implementation plan (at the proposal time) 9 regcomp.c regexec.c regparse.c Regexp matching flow (Onigmo): regcomp.rb regexec.rb written in Ruby (so JIT powered ) by exposing internals (parser, char class) compiler matching VM

Slide 10

Slide 10 text

Ruby's Regexp is not Onigmo.

Slide 11

Slide 11 text

Onigmo is a Regexp engine not only for Ruby.

Slide 12

Slide 12 text

Misperception about Ruby's Regexp 12 regcomp.c regexec.c regparse.c matching VM Onigmo Regexp matching flow (Ruby): Ruby's preprocess (2) Ruby's preprocess (1) Ruby's Regexp compiler Problem area

Slide 13

Slide 13 text

We need a "Ruby-dedicated" Regexp engine!

Slide 14

Slide 14 text

Onigmo should be made over!

Slide 15

Slide 15 text

Let's start "Onigmo remake project"! Chapter 1 -Fin-

Slide 16

Slide 16 text

Today's contents 16 Why remake Onigmo? Chapter 2 "Onigmo remake project" - Current status Chapter 3 To the Regexp + JIT dream... Chapter 4 Ruby's Regexp vs Onigmo's Chapter 1

Slide 17

Slide 17 text

Ruby's Regexp preprocessing is crazy.

Slide 18

Slide 18 text

1 pattern string requires 3 preprocessing + 1 true processing.

Slide 19

Slide 19 text

4-stage Regexp (pre)processing 19 Handling escape sequences on parsing as Ruby (prism/parse.c) Handling escape sequences before compiling Regexp (re.c) Fast checking for collecting named captures (prism/regexp.c) Parsing with Onigmo (regparse.c) 1 2 3 4

Slide 20

Slide 20 text

What each stage does (1) 20 Handling escape sequences on parsing as Ruby (prism/prism.c) 1 /(?\woo)\ \u{1F600}/ =~ "foo " p /x\ y/ # => ??? original /(?\woo)\u{1F600}/ =~ "foo " p /xy/ # => /xy/ processed (1)

Slide 21

Slide 21 text

What each stage does (2) 21 Fast checking for collecting named captures (prism/regexp.c) 2 /(?\woo)\u{1F600}/ =~ "foo " p /xy/ # => /xy/ processed (1) /(?\woo)\u{1F600}/ =~ "foo " x = $~[:x] if $~ p /xy/ # => /xy/ processed (2)

Slide 22

Slide 22 text

What each stage does (3) 22 Handling escape sequences before compiling Regexp (re.c) 3 /(?\woo)\u{1F600}/ =~ "foo " x = $~[:x] if $~ p /xy/ # => /xy/ processed (2) s = "(?\woo) " e = UTF-8 onigmo_parse(s, e) internal

Slide 23

Slide 23 text

What each stage does (4) 23 Parsing with Onigmo (regparse.c) 4 s = "(?\woo) " e = UTF-8 onigmo_parse(s, e) internal list enclose string "oo" string " " list ctype \w internal AST

Slide 24

Slide 24 text

Details of Regexp preprocessing in Ruby • These preprocessing stages are for the Prism case. • Regexp preprocessing stages depend on Ruby parser and how a Regexp value created. • parse.y also has its own preprocessing stage, but it uses Onigmo for collecting named captures. • When Regexp.new(...) is called, preprocessing is started from stage 3. 24

Slide 25

Slide 25 text

prism/prism.c prism/regexp.c re.c Prism /.../ path Regexp.new('...') path regparse.c parse.y /.../ path parse.y Ruby's Regexp regcomp.c regexec.c

Slide 26

Slide 26 text

Preprocessing chaos

Slide 27

Slide 27 text

Bugs coming from preprocessing chaos • Double-, triple-, and quadruple-escape processing is highly problematic. There are countless bugs. • e.g., 1. `/\c?/ =~ "\x7F"`, but `Regexp.new('\c?') !~ "\x7F"`. 2. `/(?<\x61>x)/ =~ "x"` raises IndexError. 3. `/[]]/` is an error in Prism, but it works in parse.y. 27

Slide 28

Slide 28 text

28 Implementation plan (at the proposal time) 9 regcomp.c regexec.c regparse.c Regexp matching flow (Onigmo): regcomp.rb regexec.rb written in Ruby (so JIT powered ) by exposing internals (parser, char class) compiler matching VM

Slide 29

Slide 29 text

Onigmo.parse ?

Slide 30

Slide 30 text

Onigmo's Regexp ≠ Ruby's Regexp

Slide 31

Slide 31 text

Onigmo.parse Regexp.parse

Slide 32

Slide 32 text

prism/prism.c prism/regexp.c re.c Prism /.../ path Regexp.new('...') path regparse.c parse.y /.../ path parse.y Ruby's Regexp regcomp.c regexec.c Regexp.parse path (?)

Slide 33

Slide 33 text

Seriously?

Slide 34

Slide 34 text

Ideal architecture 34 Prism /.../ path parse.y /.../ path Regexp.new path Regexp.parse path comp.c exec.c parse.c

Slide 35

Slide 35 text

Onigmo is a Regexp engine not only for Ruby.

Slide 36

Slide 36 text

We need a "Ruby-dedicated" Regexp engine!

Slide 37

Slide 37 text

Do you understand?

Slide 38

Slide 38 text

Do you wanna remake Onigmo? Chapter 2 -Fin-

Slide 39

Slide 39 text

Today's contents 39 Why remake Onigmo? Chapter 2 "Onigmo remake project" - Current status Chapter 3 To the Regexp + JIT dream... Chapter 4 Ruby's Regexp vs Onigmo's Chapter 1

Slide 40

Slide 40 text

Project Naraku 5IF0OJHNP3FNBLF1SPKFDU

Slide 41

Slide 41 text

Goals of Project Naraku 1.Modern & clean architecture 2.Providing the Ruby's Regexp specification 3.User friendly new features 4.Performance improvement (with JIT?) 41 Creating a "Ruby-dedicated" Regexp engine

Slide 42

Slide 42 text

Goals of Project Naraku 1.Modern & clean architecture 2.Providing the Ruby's Regexp specification 3.User friendly new features 4.Performance improvement (with JIT?) 42 Creating a "Ruby-dedicated" Regexp engine

Slide 43

Slide 43 text

History of Oniguruma, Onigmo, and Ruby 43 Oniguruma development started Ruby 1.9 2002 2007 2011 Onigmo forked from Oniguruma 2013 Ruby 2.0 0OJHVSVNB Onigmo 2022 Ruby 3.2.0 (memoization) Ruby's Regexp engine:

Slide 44

Slide 44 text

others 839 reg*.c 404 Number of `goto` statements: reg*.c (Onigmo) vs. others

Slide 45

Slide 45 text

`goto` hell

Slide 46

Slide 46 text

46 Ruby's Regexp Eliminate! Onigmo's Regexp Naraku

Slide 47

Slide 47 text

Goals of Project Naraku 1.Modern & clean architecture 2.Providing the Ruby's Regexp specification 3.User friendly new features 4.Performance improvement (with JIT?) 47 Creating a "Ruby-dedicated" Regexp engine

Slide 48

Slide 48 text

"Compatibiltiy" Myth

Slide 49

Slide 49 text

Ruby's Regexp only in our brains/recognition

Slide 50

Slide 50 text

prism/prism.c prism/regexp.c re.c Prism /.../ path Regexp.new('...') path regparse.c parse.y /.../ path parse.y Ruby's Regexp regcomp.c regexec.c Canon!

Slide 51

Slide 51 text

Goals of Project Naraku 1.Modern & clean architecture 2.Fix the Ruby's Regexp behavior 3.User friendly new features 4.Performance improvement (with JIT?) 51 Creating a "Ruby-dedicated" Regexp engine

Slide 52

Slide 52 text

`i` flag is complex due to Unicode full case folding

Slide 53

Slide 53 text

/(?i)Straße/ matches 704 variants.

Slide 54

Slide 54 text

s t r a s s e S T R A S S E ſ ſt st ſ ß ẞ ſ

Slide 55

Slide 55 text

Introduce `I` flag ASCII only case-folding

Slide 56

Slide 56 text

Goals of Project Naraku 1.Modern & clean architecture 2.Providing the Ruby's Regexp specification 3.User friendly new features 4.Performance improvement (with JIT?) 56 Creating a "Ruby-dedicated" Regexp engine

Slide 57

Slide 57 text

Current status of Project Naraku

Slide 58

Slide 58 text

https://github.com/ makenowjust/naraku

Slide 59

Slide 59 text

TODO •[-] Encoding •[x] UTF-8, Shift_JIS, ASCII-8BIT / [ ] Others •[x] Parser •[ ] Matching VM / [ ] Compilation 59

Slide 60

Slide 60 text

Parser is implemeted.

Slide 61

Slide 61 text

Ideal Current architecture 61 NarakuRuby.parse path Prism /.../ path Regexp.new path comp.c exec.c parse.c not yet implemented Ruby prototyping for JIT speed-up comp.rb exec.rb Chapter 3 -Fin-

Slide 62

Slide 62 text

Today's contents 62 Why remake Onigmo? Chapter 2 "Onigmo remake project" - Current status Chapter 3 To the Regexp + JIT dream... Chapter 4 Ruby's Regexp vs Onigmo's Chapter 1

Slide 63

Slide 63 text

Goals of Project Naraku 1.Modern & clean architecture 2.Providing the Ruby's Regexp specification 3.User friendly new features 4.Performance improvement (with JIT?) 63 Creating a "Ruby-dedicated" Regexp engine

Slide 64

Slide 64 text

NarakuRuby::DFA • An on-the-fly DFA construction Regexp engine (like Go's `regexp` package) written purely in Ruby. • Limitations: UTF-8 only. No lookarounds, backreferences, and sub-exp calls. 64

Slide 65

Slide 65 text

NarakuRuby::DFA Important disclaimer • This engine was built solely for "Regexp + [YZ]JIT" experiments. • It does NOT represent the future architectural direction of Project Naraku and Ruby. 65

Slide 66

Slide 66 text

Benchmark results (in iteration-per-second) 66 #FODIDBTF NarakuRuby::DFA.match? (no JIT) NarakuRuby::DFA.match? (w/ YJIT) YJIT / no JIT NarakuRuby::DFA.match? (w/ ZJIT) ;+*5OP+*5 MJUFSBM JT JT YGBTUFS JT YGBTUFS BMUFSOBUJPO JT JT YGBTUFS JT YGBTUFS SFQFUJUJPO HSFFEZ JT JT YGBTUFS JT YGBTUFS SFQFUJUJPO BNCJHVPVT JT JT YGBTUFS JT YGBTUFS DIBS@DMBTT JT LJT YGBTUFS JT YGBTUFS BODIPS JT LJT YGBTUFS JT YGBTUFS Comparing with [YZ]JIT

Slide 67

Slide 67 text

Benchmark results (in iteration-per-second) 67 #FODIDBTF Regexp#match? (Onigmo) NarakuRuby::DFA.match? (Naraku) /BSBLV0OJHNP MJUFSBM LJT JT YTMPXFS BMUFSOBUJPO LJT JT YTMPXFS SFQFUJUJPOHSFFEZ LJT JT YTMPXFS SFQFUJUJPOBNCJHVPVT JT JT YTMPXFS DIBS@DMBTT LJT LJT YTMPXFS BODIPS LJT LJT YTMPXFS Comparing with Onigmo (YJIT enabled)

Slide 68

Slide 68 text

Onigmo (and C) is fast.

Slide 69

Slide 69 text

Over optimization

Slide 70

Slide 70 text

Ahead-of-time (AOT) DFA construction + inline source expansion (source generation & eval)

Slide 71

Slide 71 text

No content

Slide 72

Slide 72 text

Benchmark results (in iteration-per-second) 72 #FODIDBTF Regexp#match? (Onigmo) NarakuRuby::DFA.match? (Naraku) Over-optimized NarakuRubyDFA.match? (OO Naraku) 00/BSBLV0OJHNP MJUFSBM LJT JT LJT YTMPXF SFQFUJUJPO HSFFEZ LJT JT LJT YTMPXFS SFQFUJUJPO BNCJHVPVT JT JT LJT YGBTUFS with YJIT

Slide 73

Slide 73 text

Lessons from benchmarks 73 • YJIT enables a pure-Ruby Regexp engine fast as Onigmo. • However, it causes a maintainability issue. • Perhaps, my program did not obtain the full ZJIT power. I try to learn the ZJIT architecture.

Slide 74

Slide 74 text

Next plan for Project Naraku 74 • To prevent performance degradation, we first go for the realistic way. parse.c comp.rb exec.c C side Ruby side matching VM

Slide 75

Slide 75 text

Direct byte-code compilation issue 75 concat literal "a" literal "c" alt /(a|b)c/ push :br char "a" jump :exit :br char "b" :exit char "c" byte code literal "b" What prev char? Current

Slide 76

Slide 76 text

IR (Internal Representation) for Regexp 76 concat literal "a" literal "c" alt /(a|b)c/ push :br char "a" jump :exit :br char "b" :exit char "c" byte code literal "b" push char "a" char "b" char "c" IR Future

Slide 77

Slide 77 text

Memoization and IR 77 concat literal "a" literal "b" quantifier {1,*} /a+b/ :enter char "a" push :exit jump :enter :exit char "b" Current char "a" char "b" push Future / byte code IR

Slide 78

Slide 78 text

Next next future 78 • This architecture is easy to introduce Ruby matching VM! exec.rb pure Ruby matching VM! parse.c C side Ruby side 78 exec.c matching VM comp.rb

Slide 79

Slide 79 text

Remaining tasks 79 • Matching VM / Compilation • Other encodings (GB18030, etc) • Connecting to Ruby (replacing Onigmo with Naraku) • Subset for Prism (?) • Java or WebAssembly binding for JRuby (?)

Slide 80

Slide 80 text

Made it!

Slide 81

Slide 81 text

To be continued...