# Oniguruma lookbehind question

**URL:** <https://rubytalk.org/t/oniguruma-lookbehind-question/23917>\
**Category:** ruby-talk\
**Created:** [7 January 2006 02:56 UTC](https://rubytalk.org/t/oniguruma-lookbehind-question/23917 "2006-01-07T02:56:50Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![David\_A\_Black3](https://avatars.discourse-cdn.com/v4/letter/d/6a8cbe/32.png) [@David\_A\_Black3](https://rubytalk.org/u/David_A_Black3)\
**Post date:** [7 January 2006 02:56 UTC](https://rubytalk.org/t/oniguruma-lookbehind-question/23917/1 "2006-01-07T02:56:50Z")

</div>

Hi --

For some reason, lookbehind and alternation seem not to be playing  
together in a little Oniguruma test. This is based on the string  
splitting thread from a little while ago this evening, and uses a CVS  
1.9.0 Ruby acquired about 1/2 an hour ago.

&nbsp;&nbsp;&nbsp;str = %Q{abc def "ghi jkl" mno}

&nbsp;&nbsp;&nbsp;# Look for "..." but just get the ... part:  
&nbsp;&nbsp;&nbsp;re1 = /(?\<=")[^"]+(?=")/

&nbsp;&nbsp;&nbsp;# Test that:  
&nbsp;&nbsp;&nbsp;p str.scan(re1) # =\> ["ghi jkl"]

&nbsp;&nbsp;&nbsp;# Now, do the same thing \*or\* \S+. This should, I think,  
&nbsp;&nbsp;&nbsp;# pick up the abc, def, and mno substrings too.

&nbsp;&nbsp;&nbsp;re2 = /((?\<=")[^"]+(?="))|(\S+)/

&nbsp;&nbsp;&nbsp;# But it doesn't; the part before the alternation never  
&nbsp;&nbsp;&nbsp;# matches, even though it did before (as shown by the  
&nbsp;&nbsp;&nbsp;# captures):

&nbsp;&nbsp;&nbsp;p str.scan(re2)  
&nbsp;&nbsp;&nbsp;# =\> [[nil, "abc"], [nil, "def"], [nil, "\"ghi"], [nil, "jkl\""],  
&nbsp;&nbsp;&nbsp;# [nil, "mno"]]

I know that's all a bit cluttered, but the basic thing is that a  
sub-pattern using lookbehind doesn't seem to match any more when  
there's an alternation. Instead, only the second alternative ever  
matches.

Does anyone know why?

David

> **···**
>
> --  
> David A. Black  
> [dblack@wobblini.net](mailto:dblack@wobblini.net)
> 
> "Ruby for Rails", from Manning Publications, coming April 2006!
> 
> > **[Ruby for Rails](https://www.manning.com/books/ruby-for-rails)**
> >
> > NEWER EDITION AVAILABLE
> > 
> > The Well-Grounded Rubyist, Second Edition is now available. An eBook of the previous edition, The Well-Grounded Rubyist is included at no additional cost when you buy the revised edition! 
> > 
> > 
> > Ruby for Rails helps Rails...

---

<div class="post-metadata">

**Author:** ![K.Kosako1](https://avatars.discourse-cdn.com/v4/letter/k/a5b964/32.png) [@K.Kosako1](https://rubytalk.org/u/K.Kosako1)\
**Post date:** [7 January 2006 03:28 UTC](https://rubytalk.org/t/oniguruma-lookbehind-question/23917/2 "2006-01-07T03:28:39Z")

</div>

dblack@wobblini.net wrote:

> For some reason, lookbehind and alternation seem not to be playing  
> together in a little Oniguruma test. This is based on the string  
> splitting thread from a little while ago this evening, and uses a CVS  
> 1.9.0 Ruby acquired about 1/2 an hour ago.
> 
> &nbsp;&nbsp;str = %Q{abc def "ghi jkl" mno}
> 
> &nbsp;&nbsp;# Look for "..." but just get the ... part:  
> &nbsp;&nbsp;re1 = /(?\<=")[^"]+(?=")/
> 
> &nbsp;&nbsp;# Test that:  
> &nbsp;&nbsp;p str.scan(re1) # =\> ["ghi jkl"]
> 
> &nbsp;&nbsp;# Now, do the same thing \*or\* \S+. This should, I think,  
> &nbsp;&nbsp;# pick up the abc, def, and mno substrings too.
> 
> &nbsp;&nbsp;re2 = /((?\<=")[^"]+(?="))|(\S+)/
> 
> &nbsp;&nbsp;# But it doesn't; the part before the alternation never  
> &nbsp;&nbsp;# matches, even though it did before (as shown by the  
> &nbsp;&nbsp;# captures):
> 
> &nbsp;&nbsp;p str.scan(re2)  
> &nbsp;&nbsp;# =\> [[nil, "abc"], [nil, "def"], [nil, "\"ghi"], [nil, "jkl\""],  
> &nbsp;&nbsp;# [nil, "mno"]]
> 
> I know that's all a bit cluttered, but the basic thing is that a  
> sub-pattern using lookbehind doesn't seem to match any more when  
> there's an alternation. Instead, only the second alternative ever  
> matches.

Is this pattern work for you?

str = %Q{abc def "ghi jkl" mno}  
re3 = /((?\<=")[^"]+(?="))|([\S&&[^"]]+)/  
p str.scan(re3) #=\> [[nil, "abc"], [nil, "def"], ["ghi jkl", nil], [nil, "mno"]]

> **···**
>
> --  
> K.Kosako

---

<div class="post-metadata">

**Author:** ![Ross\_Bamford2](https://avatars.discourse-cdn.com/v4/letter/r/e47774/32.png) [@Ross\_Bamford2](https://rubytalk.org/u/Ross_Bamford2)\
**Post date:** [7 January 2006 03:43 UTC](https://rubytalk.org/t/oniguruma-lookbehind-question/23917/3 "2006-01-07T03:43:00Z")

</div>

I'm not at all sure about this, but this is my take on it. Firstly, is this the behaviour you expected?

&nbsp;&nbsp;str = %Q{abc def "ghi jkl" mno}  
&nbsp;&nbsp;re2 = /(?:(?\<=")[^"]+(?="))|\S+/

&nbsp;&nbsp;p str.scan(re2)  
&nbsp;&nbsp;# =\> ["abc", "def", "\"ghi", "jkl\"", "mno"]

?

If so, then I believe the problem is something to do with the fact that lookaround is atomic, so that when used with capturing groups and alternations you sometimes experience problems because the regex immediately forgets the (zero-width, remember) lookaround match, so that by the time it comes to that 'or' it doesn't have the information to compare.

Generally, there are restrictions with lookaround (esp lookbehind) matching, and especially when matching regexps. So far my experiments with Oniguruma suggest it's fairly sophisticated in this respect, supporting stuff like varying-width alternations, fixed repetition and optional groups in lookbehind, but of course still no star and plus.

Anyway, that's what I think. Hope it helps 🙂

Cheers,

> **···**
>
> On Sat, 07 Jan 2006 02:56:50 -0000, \<dblack@wobblini.net\> wrote:
> 
> > Hi --
> > 
> > For some reason, lookbehind and alternation seem not to be playing  
> > together in a little Oniguruma test. This is based on the string  
> > splitting thread from a little while ago this evening, and uses a CVS  
> > 1.9.0 Ruby acquired about 1/2 an hour ago.
> > 
> > &nbsp;&nbsp;&nbsp;str = %Q{abc def "ghi jkl" mno}
> > 
> > &nbsp;&nbsp;&nbsp;# Look for "..." but just get the ... part:  
> > &nbsp;&nbsp;&nbsp;re1 = /(?\<=")[^"]+(?=")/
> > 
> > &nbsp;&nbsp;&nbsp;# Test that:  
> > &nbsp;&nbsp;&nbsp;p str.scan(re1) # =\> ["ghi jkl"]
> > 
> > &nbsp;&nbsp;&nbsp;# Now, do the same thing \*or\* \S+. This should, I think,  
> > &nbsp;&nbsp;&nbsp;# pick up the abc, def, and mno substrings too.
> > 
> > &nbsp;&nbsp;&nbsp;re2 = /((?\<=")[^"]+(?="))|(\S+)/
> > 
> > &nbsp;&nbsp;&nbsp;# But it doesn't; the part before the alternation never  
> > &nbsp;&nbsp;&nbsp;# matches, even though it did before (as shown by the  
> > &nbsp;&nbsp;&nbsp;# captures):
> > 
> > &nbsp;&nbsp;&nbsp;p str.scan(re2)  
> > &nbsp;&nbsp;&nbsp;# =\> [[nil, "abc"], [nil, "def"], [nil, "\"ghi"], [nil, "jkl\""],  
> > &nbsp;&nbsp;&nbsp;# [nil, "mno"]]
> > 
> > I know that's all a bit cluttered, but the basic thing is that a  
> > sub-pattern using lookbehind doesn't seem to match any more when  
> > there's an alternation. Instead, only the second alternative ever  
> > matches.
> > 
> > Does anyone know why?
> 
> --  
> Ross Bamford - rosco@roscopeco.remove.co.uk

---

<div class="post-metadata">

**Author:** ![Xavier\_Noria](https://yyz1.discourse-cdn.com/flex029/user_avatar/rubytalk.org/xavier_noria/32/1867_2.png) [@Xavier\_Noria](https://rubytalk.org/u/Xavier_Noria)\
**Post date:** [7 January 2006 03:52 UTC](https://rubytalk.org/t/oniguruma-lookbehind-question/23917/4 "2006-01-07T03:52:49Z")

</div>

It shouldn't, since pattern-matching goes left-to-right \S will match the quote before the first half of the regexp gets a chance, since it wants to match the first character \_after\_ the quote.

-- fxn

> **···**
>
> On Jan 7, 2006, at 3:56, dblack@wobblini.net wrote:
> 
> > &nbsp;&nbsp;# Now, do the same thing \*or\* \S+. This should, I think,  
> > &nbsp;&nbsp;# pick up the abc, def, and mno substrings too.
> > 
> > &nbsp;&nbsp;re2 = /((?\<=")[^"]+(?="))|(\S+)/
> > 
> > &nbsp;&nbsp;# But it doesn't; the part before the alternation never  
> > &nbsp;&nbsp;# matches,

---

<div class="post-metadata">

**Author:** ![Ross\_Bamford2](https://avatars.discourse-cdn.com/v4/letter/r/e47774/32.png) [@Ross\_Bamford2](https://rubytalk.org/u/Ross_Bamford2)\
**Post date:** [7 January 2006 03:47 UTC](https://rubytalk.org/t/oniguruma-lookbehind-question/23917/5 "2006-01-07T03:47:59Z")

</div>

Oops, I mean nested regexps.

> **···**
>
> On Sat, 07 Jan 2006 03:36:34 -0000, Ross Bamford \<rosco@roscopeco.remove.co.uk\> wrote:
> 
> > Generally, there are restrictions with lookaround (esp lookbehind) matching, and especially when matching regexps.
> 
> --  
> Ross Bamford - rosco@roscopeco.remove.co.uk

---

<div class="post-metadata">

**Author:** ![David\_A\_Black3](https://avatars.discourse-cdn.com/v4/letter/d/6a8cbe/32.png) [@David\_A\_Black3](https://rubytalk.org/u/David_A_Black3)\
**Post date:** [7 January 2006 03:49 UTC](https://rubytalk.org/t/oniguruma-lookbehind-question/23917/6 "2006-01-07T03:49:32Z")

</div>

Hi --

> > For some reason, lookbehind and alternation seem not to be playing  
> > together in a little Oniguruma test. This is based on the string  
> > splitting thread from a little while ago this evening, and uses a CVS  
> > 1.9.0 Ruby acquired about 1/2 an hour ago.
> > 
> > &nbsp;&nbsp;str = %Q{abc def "ghi jkl" mno}
> > 
> > &nbsp;&nbsp;# Look for "..." but just get the ... part:  
> > &nbsp;&nbsp;re1 = /(?\<=")[^"]+(?=")/
> > 
> > &nbsp;&nbsp;# Test that:  
> > &nbsp;&nbsp;p str.scan(re1) # =\> ["ghi jkl"]
> > 
> > &nbsp;&nbsp;# Now, do the same thing \*or\* \S+. This should, I think,  
> > &nbsp;&nbsp;# pick up the abc, def, and mno substrings too.
> > 
> > &nbsp;&nbsp;re2 = /((?\<=")[^"]+(?="))|(\S+)/
> > 
> > &nbsp;&nbsp;# But it doesn't; the part before the alternation never  
> > &nbsp;&nbsp;# matches, even though it did before (as shown by the  
> > &nbsp;&nbsp;# captures):
> > 
> > &nbsp;&nbsp;p str.scan(re2)  
> > &nbsp;&nbsp;# =\> [[nil, "abc"], [nil, "def"], [nil, "\"ghi"], [nil, "jkl\""],  
> > &nbsp;&nbsp;# [nil, "mno"]]
> > 
> > I know that's all a bit cluttered, but the basic thing is that a  
> > sub-pattern using lookbehind doesn't seem to match any more when  
> > there's an alternation. Instead, only the second alternative ever  
> > matches.
> 
> Is this pattern work for you?
> 
> str = %Q{abc def "ghi jkl" mno}  
> re3 = /((?\<=")[^"]+(?="))|([\S&&[^"]]+)/  
> p str.scan(re3) #=\> [[nil, "abc"], [nil, "def"], ["ghi jkl", nil], [nil, "mno"]]

No; I get this:

[[nil, "\""], ["ghi jkl", nil], [nil, "\""]]

This:

&nbsp;&nbsp;&nbsp;/(?\<=")[^"]+(?=")|[^\s"]+/

gives me the result you got for your re3. But it still seems to be  
based on the right-hand alternate being checked first.

David

> **···**
>
> On Sat, 7 Jan 2006, K.Kosako wrote:
> 
> > dblack@wobblini.net wrote:
> 
> --  
> David A. Black  
> dblack@wobblini.net
> 
> "Ruby for Rails", from Manning Publications, coming April 2006!
> 
> > **[Ruby for Rails](https://www.manning.com/books/ruby-for-rails)**
> >
> > NEWER EDITION AVAILABLE
> > 
> > The Well-Grounded Rubyist, Second Edition is now available. An eBook of the previous edition, The Well-Grounded Rubyist is included at no additional cost when you buy the revised edition! 
> > 
> > 
> > Ruby for Rails helps Rails...

---

<div class="post-metadata">

**Author:** ![Ross\_Bamford2](https://avatars.discourse-cdn.com/v4/letter/r/e47774/32.png) [@Ross\_Bamford2](https://rubytalk.org/u/Ross_Bamford2)\
**Post date:** [7 January 2006 04:08 UTC](https://rubytalk.org/t/oniguruma-lookbehind-question/23917/7 "2006-01-07T04:08:00Z")

</div>

Argh, I meant 'if not'. I'm too tired now, I've spent too long perfecting this quiz thing...  
I was guessing it was dropping the match without the capture and so always matching the alternate but that's not right.

Sorry.

> **···**
>
> On Sat, 07 Jan 2006 03:36:34 -0000, Ross Bamford \<rosco@roscopeco.remove.co.uk\> wrote:
> 
> > If so, then I believe the problem is something to do with the fact that lookaround is atomic
> 
> --  
> Ross Bamford - rosco@roscopeco.remove.co.uk

---

<div class="post-metadata">

**Author:** ![David\_A\_Black3](https://avatars.discourse-cdn.com/v4/letter/d/6a8cbe/32.png) [@David\_A\_Black3](https://rubytalk.org/u/David_A_Black3)\
**Post date:** [7 January 2006 04:10 UTC](https://rubytalk.org/t/oniguruma-lookbehind-question/23917/8 "2006-01-07T04:10:04Z")

</div>

Hi --

> **···**
>
> On Sat, 7 Jan 2006, Xavier Noria wrote:
> 
> > On Jan 7, 2006, at 3:56, dblack@wobblini.net wrote:
> > 
> > > # Now, do the same thing \*or\* \S+. This should, I think,  
> > > # pick up the abc, def, and mno substrings too.
> > > 
> > > re2 = /((?\<=")[^"]+(?="))|(\S+)/
> > > 
> > > # But it doesn't; the part before the alternation never  
> > > # matches,
> > 
> > It shouldn't, since pattern-matching goes left-to-right \S will match the quote before the first half of the regexp gets a chance, since it wants to match the first character \_after\_ the quote.
> 
> OK, I see. I was somehow discounting the fact that the first "  
> \*itself\* doesn't match the left-hand alternate.
> 
> Thanks --
> 
> David
> 
> --  
> David A. Black  
> dblack@wobblini.net
> 
> "Ruby for Rails", from Manning Publications, coming April 2006!
> 
> > **[Ruby for Rails](https://www.manning.com/books/ruby-for-rails)**
> >
> > NEWER EDITION AVAILABLE
> > 
> > The Well-Grounded Rubyist, Second Edition is now available. An eBook of the previous edition, The Well-Grounded Rubyist is included at no additional cost when you buy the revised edition! 
> > 
> > 
> > Ruby for Rails helps Rails...

---

<div class="post-metadata">

**Author:** ![David\_A\_Black3](https://avatars.discourse-cdn.com/v4/letter/d/6a8cbe/32.png) [@David\_A\_Black3](https://rubytalk.org/u/David_A_Black3)\
**Post date:** [7 January 2006 04:17 UTC](https://rubytalk.org/t/oniguruma-lookbehind-question/23917/9 "2006-01-07T04:17:35Z")

</div>

Hi --

> **···**
>
> On Sat, 7 Jan 2006, K.Kosako wrote:
> 
> > Is this pattern work for you?
> > 
> > str = %Q{abc def "ghi jkl" mno}  
> > re3 = /((?\<=")[^"]+(?="))|([\S&&[^"]]+)/  
> > p str.scan(re3) #=\> [[nil, "abc"], [nil, "def"], ["ghi jkl", nil], [nil, "mno"]]
> 
> Sorry -- I tested that accidentally with an old 1.9.0. Yes, with  
> today's I do get the same as you.
> 
> (And I also understand why it's choosing the right-hand alternate 🙂  
> See later posts in thread.)
> 
> David
> 
> --  
> David A. Black  
> dblack@wobblini.net
> 
> "Ruby for Rails", from Manning Publications, coming April 2006!
> 
> > **[Ruby for Rails](https://www.manning.com/books/ruby-for-rails)**
> >
> > NEWER EDITION AVAILABLE
> > 
> > The Well-Grounded Rubyist, Second Edition is now available. An eBook of the previous edition, The Well-Grounded Rubyist is included at no additional cost when you buy the revised edition! 
> > 
> > 
> > Ruby for Rails helps Rails...

---

<div class="post-metadata">

**Author:** ![David\_A\_Black3](https://avatars.discourse-cdn.com/v4/letter/d/6a8cbe/32.png) [@David\_A\_Black3](https://rubytalk.org/u/David_A_Black3)\
**Post date:** [7 January 2006 04:12 UTC](https://rubytalk.org/t/oniguruma-lookbehind-question/23917/10 "2006-01-07T04:12:53Z")

</div>

Hi --

> > If so, then I believe the problem is something to do with the fact that lookaround is atomic
> 
> Argh, I meant 'if not'. I'm too tired now, I've spent too long perfecting this quiz thing...  
> I was guessing it was dropping the match without the capture and so always matching the alternate but that's not right.

See Xavier's post. My mistake was, essentially, expecting the first "  
to "know" that it was supposed to match a zero-width condition  
governing the state one character later. Instead, of course, it  
asserts itself as a character in its own right; fails to match the  
first alternate; and does match the second.

So [^\s"]+ is indeed probably the best thing. (Other than the  
appropriate real libraries, of course 🙂

David

> **···**
>
> On Sat, 7 Jan 2006, Ross Bamford wrote:
> 
> > On Sat, 07 Jan 2006 03:36:34 -0000, Ross Bamford \> \<rosco@roscopeco.remove.co.uk\> wrote:
> 
> --  
> David A. Black  
> dblack@wobblini.net
> 
> "Ruby for Rails", from Manning Publications, coming April 2006!
> 
> > **[Ruby for Rails](https://www.manning.com/books/ruby-for-rails)**
> >
> > NEWER EDITION AVAILABLE
> > 
> > The Well-Grounded Rubyist, Second Edition is now available. An eBook of the previous edition, The Well-Grounded Rubyist is included at no additional cost when you buy the revised edition! 
> > 
> > 
> > Ruby for Rails helps Rails...

---

<div class="post-metadata">

**Author:** ![Ross\_Bamford2](https://avatars.discourse-cdn.com/v4/letter/r/e47774/32.png) [@Ross\_Bamford2](https://rubytalk.org/u/Ross_Bamford2)\
**Post date:** [7 January 2006 04:28 UTC](https://rubytalk.org/t/oniguruma-lookbehind-question/23917/11 "2006-01-07T04:28:00Z")

</div>

Oh Damn it, yeah I see now. Wish I'd held my tongue now 😃

> **···**
>
> On Sat, 07 Jan 2006 04:12:53 -0000, \<dblack@wobblini.net\> wrote:
> 
> > Hi --
> > 
> > On Sat, 7 Jan 2006, Ross Bamford wrote:
> > 
> > > On Sat, 07 Jan 2006 03:36:34 -0000, Ross Bamford \>\> \<rosco@roscopeco.remove.co.uk\> wrote:
> > > 
> > > > If so, then I believe the problem is something to do with the fact that lookaround is atomic
> > > 
> > > Argh, I meant 'if not'. I'm too tired now, I've spent too long perfecting this quiz thing...  
> > > I was guessing it was dropping the match without the capture and so always matching the alternate but that's not right.
> > 
> > See Xavier's post. My mistake was, essentially, expecting the first "  
> > to "know" that it was supposed to match a zero-width condition  
> > governing the state one character later. Instead, of course, it  
> > asserts itself as a character in its own right; fails to match the  
> > first alternate; and does match the second.
> 
> --  
> Ross Bamford - rosco@roscopeco.remove.co.uk
