# \[Q\] specify start postion of Regexp matching

**URL:** <https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487>\
**Category:** ruby-talk\
**Created:** [25 November 2007 15:20 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487 "2007-11-25T15:20:25Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![Makoto\_Kuwata](https://avatars.discourse-cdn.com/v4/letter/m/a4c791/32.png) [@Makoto\_Kuwata](https://rubytalk.org/u/Makoto_Kuwata)\
**Post date:** [25 November 2007 15:20 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/1 "2007-11-25T15:20:25Z")

</div>

Hi, all.

Is it possible to specify start position of Regexp matching?

&nbsp;&nbsp;&nbsp;&nbsp;str = "foo bar baz"  
&nbsp;&nbsp;&nbsp;&nbsp;m = /ba/.match(str)  
&nbsp;&nbsp;&nbsp;&nbsp;p m.begin(0) #=\> 4  
&nbsp;&nbsp;&nbsp;&nbsp;m = /ba/.match(str, 5) # is it possible?  
&nbsp;&nbsp;&nbsp;&nbsp;p m.begin(0) #=\> 8 (if possible)

If it is possible, some kind of parser or scanner can be  
implemented easily.  
# StringScanner is a litte too big, I think.

> **···**
>
> --  
> makoto kuwata

---

<div class="post-metadata">

**Author:** ![Eric\_I2](https://avatars.discourse-cdn.com/v4/letter/e/22d042/32.png) [@Eric\_I2](https://rubytalk.org/u/Eric_I2)\
**Post date:** [25 November 2007 15:39 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/2 "2007-11-25T15:39:59Z")

</div>

You could try something like this:

&nbsp;&nbsp;&nbsp;&nbsp;m = /^.{5,}(ba)/.match(str)  
&nbsp;&nbsp;&nbsp;&nbsp;p m.begin(1)

In the regular expression, you're saying start at the beginning and  
skip at least 5 characters. But then we have to use parens to "note"  
the part you're interested in, and then we have to pass 1 rather than  
0 to begin, so it reports the location of the first noted match (0  
would report where the entire Regexp matched, and that would be the  
beginning of the line).

An alternative would be to slice the first n characters off the front  
of the string and then do the match.

Eric

> **···**
>
> On Nov 25, 10:18 am, makoto kuwata \<k...@kuwata-lab.com\> wrote:
> 
> > Hi, all.
> > 
> > Is it possible to specify start position of Regexp matching?
> > 
> > &nbsp;&nbsp;&nbsp;&nbsp;str = "foo bar baz"  
> > &nbsp;&nbsp;&nbsp;&nbsp;m = /ba/.match(str)  
> > &nbsp;&nbsp;&nbsp;&nbsp;p m.begin(0) #=\> 4  
> > &nbsp;&nbsp;&nbsp;&nbsp;m = /ba/.match(str, 5) # is it possible?  
> > &nbsp;&nbsp;&nbsp;&nbsp;p m.begin(0) #=\> 8 (if possible)
> > 
> > If it is possible, some kind of parser or scanner can be  
> > implemented easily.  
> > # StringScanner is a litte too big, I think.
> 
> ====
> 
> Interested in hands-on, on-site Ruby training? See [http://LearnRuby.com](http://LearnRuby.com)  
> for information about a well-reviewed class.

---

<div class="post-metadata">

**Author:** ![Axel\_Etzold](https://avatars.discourse-cdn.com/v4/letter/a/ba8739/32.png) [@Axel\_Etzold](https://rubytalk.org/u/Axel_Etzold)\
**Post date:** [25 November 2007 16:46 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/3 "2007-11-25T16:46:39Z")

</div>

-------- Original-Nachricht --------

> Datum: Mon, 26 Nov 2007 00:20:25 +0900  
> Von: makoto kuwata \<kwa@kuwata-lab.com\>  
> An: ruby-talk@ruby-lang.org  
> Betreff: [Q] specify start postion of Regexp matching

> Hi, all.
> 
> Is it possible to specify start position of Regexp matching?
> 
> &nbsp;&nbsp;&nbsp;&nbsp;str = "foo bar baz"  
> &nbsp;&nbsp;&nbsp;&nbsp;m = /ba/.match(str)  
> &nbsp;&nbsp;&nbsp;&nbsp;p m.begin(0) #=\> 4  
> &nbsp;&nbsp;&nbsp;&nbsp;m = /ba/.match(str, 5) # is it possible?  
> &nbsp;&nbsp;&nbsp;&nbsp;p m.begin(0) #=\> 8 (if possible)
> 
> If it is possible, some kind of parser or scanner can be  
> implemented easily.  
> # StringScanner is a litte too big, I think.
> 
> --  
> makoto kuwata

Dear Makoto,

what about :

class Regexp  
&nbsp;&nbsp;def match\_index\_offset(string,start\_pos)  
&nbsp;&nbsp;&nbsp;&nbsp;temp=string[start\_pos..-1]  
&nbsp;&nbsp;&nbsp;&nbsp;ref=self.match(temp)  
&nbsp;&nbsp;&nbsp;&nbsp;return temp.index(ref[0])+start\_pos  
&nbsp;&nbsp;end  
end

str = "foo bar baz"  
m = /ba/.match\_index\_offset(str,5)  
p m

Best regards,

Axel

> **···**
>
> --  
> Psssst! Schon vom neuen GMX MultiMessenger gehört?  
> Der kann`s mit allen: [http://www.gmx.net/de/go/multimessenger](http://www.gmx.net/de/go/multimessenger)

---

<div class="post-metadata">

**Author:** ![Yukihiro\_Matsumoto2](https://avatars.discourse-cdn.com/v4/letter/y/b9e5f3/32.png) [@Yukihiro\_Matsumoto2](https://rubytalk.org/u/Yukihiro_Matsumoto2)\
**Post date:** [25 November 2007 23:24 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/4 "2007-11-25T23:24:04Z")

</div>

Hi,

> **···**
>
> In message "Re: [Q] specify start postion of Regexp matching" on Mon, 26 Nov 2007 00:20:25 +0900, makoto kuwata \<kwa@kuwata-lab.com\> writes:
> 
> > Is it possible to specify start position of Regexp matching?
> > 
> > &nbsp;&nbsp;&nbsp;str = "foo bar baz"  
> > &nbsp;&nbsp;&nbsp;m = /ba/.match(str)  
> > &nbsp;&nbsp;&nbsp;p m.begin(0) #=\> 4  
> > &nbsp;&nbsp;&nbsp;m = /ba/.match(str, 5) # is it possible?  
> > &nbsp;&nbsp;&nbsp;p m.begin(0) #=\> 8 (if possible)
> 
> str.index(/ba/, 5) ?
> 
> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;matz.

---

<div class="post-metadata">

**Author:** ![Robert\_K1](https://yyz1.discourse-cdn.com/flex029/user_avatar/rubytalk.org/robert_k1/32/1830_2.png) [@Robert\_K1](https://rubytalk.org/u/Robert_K1)\
**Post date:** [25 November 2007 16:25 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/5 "2007-11-25T16:25:05Z")

</div>

Another alternative is to use String#scan - we would have to know what the OP really wants to parse though to decide whether it's a feasible solution.

Kind regards

&nbsp;&nbsp;robert

> **···**
>
> On 25.11.2007 16:39, Eric I. wrote:
> 
> > On Nov 25, 10:18 am, makoto kuwata \<k...@kuwata-lab.com\> wrote:
> > 
> > > Hi, all.
> > > 
> > > Is it possible to specify start position of Regexp matching?
> > > 
> > > &nbsp;&nbsp;&nbsp;&nbsp;str = "foo bar baz"  
> > > &nbsp;&nbsp;&nbsp;&nbsp;m = /ba/.match(str)  
> > > &nbsp;&nbsp;&nbsp;&nbsp;p m.begin(0) #=\> 4  
> > > &nbsp;&nbsp;&nbsp;&nbsp;m = /ba/.match(str, 5) # is it possible?  
> > > &nbsp;&nbsp;&nbsp;&nbsp;p m.begin(0) #=\> 8 (if possible)
> > > 
> > > If it is possible, some kind of parser or scanner can be  
> > > implemented easily.  
> > > # StringScanner is a litte too big, I think.
> > 
> > You could try something like this:
> > 
> > &nbsp;&nbsp;&nbsp;&nbsp;m = /^.{5,}(ba)/.match(str)  
> > &nbsp;&nbsp;&nbsp;&nbsp;p m.begin(1)
> > 
> > In the regular expression, you're saying start at the beginning and  
> > skip at least 5 characters. But then we have to use parens to "note"  
> > the part you're interested in, and then we have to pass 1 rather than  
> > 0 to begin, so it reports the location of the first noted match (0  
> > would report where the entire Regexp matched, and that would be the  
> > beginning of the line).
> > 
> > An alternative would be to slice the first n characters off the front  
> > of the string and then do the match.

---

<div class="post-metadata">

**Author:** ![Makoto\_Kuwata](https://avatars.discourse-cdn.com/v4/letter/m/a4c791/32.png) [@Makoto\_Kuwata](https://rubytalk.org/u/Makoto_Kuwata)\
**Post date:** [25 November 2007 18:00 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/6 "2007-11-25T18:00:02Z")

</div>

Thank you, all.

Eric l wrote:

> You could try something like this:  
> &nbsp;&nbsp;&nbsp;&nbsp;m = /^.{5,}(ba)/.match(str)  
> &nbsp;&nbsp;&nbsp;&nbsp;p m.begin(1)

In my program, start position is variable such as  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;def f(n)  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;m = /^.{n,}(ba)/.match(str)  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;...  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;end  
In this case, /^.{n,}(ba)/ is created for each time.  
It is not effective.

Robert Klemme wrote:

> Another alternative is to use String#scan -

String#scan is useful only when regexp pattern is fixed.  
&nbsp;&nbsp;&nbsp;&nbsp;input.scan(/FIXED-REGEXP/) do ... end  
Using String#scan, it is not able to change regexp pattern  
in the loop.

Axel Etzold wrote:

> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;temp=string[start\_pos..-1]  
> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ref=self.match(temp)  
> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return temp.index(ref[0])+start\_pos

In this solution, temp substring is created every time.  
If input string is long, it is not efficient.

Thanks to all your advices.  
I'm going to propose to support start position in Regexp#match().

> **···**
>
> --  
> makoto kuwata

---

<div class="post-metadata">

**Author:** ![Makoto\_Kuwata](https://avatars.discourse-cdn.com/v4/letter/m/a4c791/32.png) [@Makoto\_Kuwata](https://rubytalk.org/u/Makoto_Kuwata)\
**Post date:** [25 November 2007 23:49 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/7 "2007-11-25T23:49:59Z")

</div>

No, String#index returns Fixnum (position), but I want MatchData.

Regexp#match(string, start=0) in Ruby1.9 is the best solution I want.  
Is there any plan to implement it into Ruby1.8?

> **···**
>
> Yukihiro Matsumoto \<m...@ruby-lang.org\> wrote:
> 
> > str.index(/ba/, 5) ?
> 
> --  
> makoto kuwata

---

<div class="post-metadata">

**Author:** ![Robert\_K1](https://yyz1.discourse-cdn.com/flex029/user_avatar/rubytalk.org/robert_k1/32/1830_2.png) [@Robert\_K1](https://rubytalk.org/u/Robert_K1)\
**Post date:** [25 November 2007 19:20 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/8 "2007-11-25T19:20:00Z")

</div>

> Robert Klemme wrote:
> 
> > Another alternative is to use String#scan -
> 
> String#scan is useful only when regexp pattern is fixed.  
> &nbsp;&nbsp;&nbsp;&nbsp;input.scan(/FIXED-REGEXP/) do ... end  
> Using String#scan, it is not able to change regexp pattern  
> in the loop.

But in various situations it is possible to use a unified regexp for scanning or a regexp that comprises all other patterns.

> Axel Etzold wrote:
> 
> > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;temp=string[start\_pos..-1]  
> > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ref=self.match(temp)  
> > &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return temp.index(ref[0])+start\_pos
> 
> In this solution, temp substring is created every time.  
> If input string is long, it is not efficient.

This is not true. Creating a substring is fairly cheap because the character buffer is not copied (copy on write).

> I'm going to propose to support start position in Regexp#match().

For the time being it's faster to use one of the other alternatives. Also, with the new regexp engine in 1.9 your feature might be present already.

Kind regards

&nbsp;&nbsp;robert

> **···**
>
> On 25.11.2007 18:58, makoto kuwata wrote:

---

<div class="post-metadata">

**Author:** ![Makoto\_Kuwata](https://avatars.discourse-cdn.com/v4/letter/m/a4c791/32.png) [@Makoto\_Kuwata](https://rubytalk.org/u/Makoto_Kuwata)\
**Post date:** [25 November 2007 23:50 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/9 "2007-11-25T23:50:00Z")

</div>

I found that it is able to get MatchData by Regexp.last\_match()  
after String#index().  
Well, I think Regexp#match(string, start=0) is the natural way,  
but String#index(regexp, start) can be the good solution.

Thank you, Matz.

> **···**
>
> makoto kuwata \<k...@kuwata-lab.com\> wrote:
> 
> > \> str.index(/ba/, 5) ?
> > 
> > No, String#index returns Fixnum (position), but I want MatchData.
> 
> --  
> makoto kuwata

---

<div class="post-metadata">

**Author:** ![Makoto\_Kuwata](https://avatars.discourse-cdn.com/v4/letter/m/a4c791/32.png) [@Makoto\_Kuwata](https://rubytalk.org/u/Makoto_Kuwata)\
**Post date:** [25 November 2007 23:25 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/10 "2007-11-25T23:25:05Z")

</div>

> \> In this solution, temp substring is created every time.  
> \> If input string is long, it is not efficient.
> 
> This is not true. Creating a substring is fairly cheap because the  
> character buffer is not copied (copy on write).

You are right. If input string is not modified, creating substring  
doesn't copy anything.  
Creating substring may be the solution I wanted.

> \> I'm going to propose to support start position in Regexp#match().
> 
> For the time being it's faster to use one of the other alternatives.  
> Also, with the new regexp engine in 1.9 your feature might be present  
> already.

I found that Regexp#match() can take optional 2nd argument which  
specifies matching start position in Ruby1.9. Good news.

Thank you, Robert.

> **···**
>
> Robert Klemme \<shortcut...@googlemail.com\> wrote:
> 
> --  
> makoto kuwata

---

<div class="post-metadata">

**Author:** ![Jordan\_Callicoat](https://avatars.discourse-cdn.com/v4/letter/j/9de053/32.png) [@Jordan\_Callicoat](https://rubytalk.org/u/Jordan_Callicoat)\
**Post date:** [26 November 2007 05:20 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/11 "2007-11-26T05:20:10Z")

</div>

What's the difference between 1.9 Regexp#match(string, start=n) and  
1.8 Regexp#match(string[n..-1])?? You have to create a sub-string with  
the 1.8 version, but according to Robert Klemme (above) it's just  
creating a pointer into the original string if you're not changing the  
substring or original string. Besides, even if you did get a copy,  
it's anonymous and should be garbage collected soon. If I understand  
everything correctly, the 1.9 version would just basically be a  
convenience feature over the 1.8 way?

$ irb19  
irb(main):001:0\> RUBY\_VERSION  
=\> "1.9.0"  
irb(main):002:0\> m = /oo/.match("foo", start=1)  
=\> #\<MatchData "oo"\>  
irb(main):003:0\> m[0]  
=\> "oo"

$ irb  
irb(main):001:0\> RUBY\_VERSION  
=\> "1.8.6"  
irb(main):002:0\> m = /oo/.match("foo"[1..-1])  
=\> #\<MatchData:0xb78777a8\>  
irb(main):003:0\> m[0]  
=\> "oo"

Regards,  
Jordan

---

<div class="post-metadata">

**Author:** ![7stud](https://avatars.discourse-cdn.com/v4/letter/7/57b2e6/32.png) [@7stud](https://rubytalk.org/u/7stud)\
**Post date:** [26 November 2007 06:07 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/12 "2007-11-26T06:07:40Z")

</div>

Jordan Callicoat wrote:

> You have to create a sub-string with  
> the 1.8 version, but according to Robert Klemme (above) it's just  
> creating a pointer into the original string if you're not changing the  
> substring or original string.

I'm having a hard time confirming that:

str = "hello"  
sub\_str = str[1, 2]

puts str.object\_id  
--\>76750

puts sub\_str.object\_id  
--\>76740

puts sub\_str.class  
--\>String

> **···**
>
> --  
> Posted via [http://www.ruby-forum.com/\](http://www.ruby-forum.com/%5C).

---

<div class="post-metadata">

**Author:** ![Jordan\_Callicoat](https://avatars.discourse-cdn.com/v4/letter/j/9de053/32.png) [@Jordan\_Callicoat](https://rubytalk.org/u/Jordan_Callicoat)\
**Post date:** [26 November 2007 06:50 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/13 "2007-11-26T06:50:00Z")

</div>

I'm not sure how to confirm it, other than just looking at the source,  
and since I'm very poor at C programming, it probably wouldn't help  
for me to try that. I'm sure Robert can demonstrate. But I will say  
that I'm not suprised that they have different object\_id, because they  
are different objects. The copy on write is just a back-end  
optimization where you pretend that two objects that point to the same  
data are unique copies in the front-end, but you don't actually move  
any data in the back-end until you have to (i,e., when one of the  
objects is changed).

Regards,  
Jordan

> **···**
>
> On Nov 26, 12:07 am, 7stud -- \<bbxx789\_0...@yahoo.com\> wrote:
> 
> > Jordan Callicoat wrote:  
> > \> You have to create a sub-string with  
> > \> the 1.8 version, but according to Robert Klemme (above) it's just  
> > \> creating a pointer into the original string if you're not changing the  
> > \> substring or original string.
> > 
> > I'm having a hard time confirming that:

---

<div class="post-metadata">

**Author:** ![George3](https://avatars.discourse-cdn.com/v4/letter/g/51bf81/32.png) [@George3](https://rubytalk.org/u/George3)\
**Post date:** [27 November 2007 05:07 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/14 "2007-11-27T05:07:39Z")

</div>

A new ruby object is created, but the string buffer that it points to  
is only copied on write.

> **···**
>
> On Nov 26, 2007 5:07 PM, 7stud -- \<bbxx789\_05ss@yahoo.com\> wrote:
> 
> > Jordan Callicoat wrote:  
> > \> You have to create a sub-string with  
> > \> the 1.8 version, but according to Robert Klemme (above) it's just  
> > \> creating a pointer into the original string if you're not changing the  
> > \> substring or original string.
> > 
> > I'm having a hard time confirming that:
> > 
> > str = "hello"  
> > sub\_str = str[1, 2]
> > 
> > puts str.object\_id  
> > --\>76750
> > 
> > puts sub\_str.object\_id  
> > --\>76740
> > 
> > puts sub\_str.class  
> > --\>String

---

<div class="post-metadata">

**Author:** ![Jordan\_Callicoat](https://avatars.discourse-cdn.com/v4/letter/j/9de053/32.png) [@Jordan\_Callicoat](https://rubytalk.org/u/Jordan_Callicoat)\
**Post date:** [26 November 2007 07:10 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/15 "2007-11-26T07:10:00Z")

</div>

Well, I did anyhow...

[http://svn.ruby-lang.org/repos/ruby/branches/ruby\_1\_8/ruby.h](http://svn.ruby-lang.org/repos/ruby/branches/ruby_1_8/ruby.h)  
[http://svn.ruby-lang.org/repos/ruby/branches/ruby\_1\_8/string.c](http://svn.ruby-lang.org/repos/ruby/branches/ruby_1_8/string.c)

And I think the functions of interest are str\_new3 and str\_new4  
(called from rb\_str\_substr). Specifically, the assignment of  
RSTRING(str2)-\>aux.shared. But like I said, I'm not great with C, so I  
could be mistaken.

Regards,  
Jordan

> **···**
>
> On Nov 26, 12:45 am, MonkeeSage \<MonkeeS...@gmail.com\> wrote:
> 
> > I'm not sure how to confirm it, other than just looking at the source,  
> > and since I'm very poor at C programming, it probably wouldn't help  
> > for me to try that.

---

<div class="post-metadata">

**Author:** ![Jordan\_Callicoat](https://avatars.discourse-cdn.com/v4/letter/j/9de053/32.png) [@Jordan\_Callicoat](https://rubytalk.org/u/Jordan_Callicoat)\
**Post date:** [26 November 2007 09:05 UTC](https://rubytalk.org/t/q-specify-start-postion-of-regexp-matching/42487/16 "2007-11-26T09:05:04Z")

</div>

Here's a test to show that my reading of the source, and Robert's  
assertion, is correct (there is probably a better way to do this...):

#!/usr/bin/env ruby

# disable GC to get fair reading of actual allocation cost  
GC.disable

def free\_megs  
&nbsp;&nbsp;&nbsp;(`free -o`.split("\n")[1].split(' ')[3].to\_i/1024).to\_s  
end

puts "Free megabytes " + free\_megs  
# make a one megabyte string  
s1 = "a" \* 1048576  
s100 = "" # placeholder to be filled in below  
# make 100 substrings of it  
0.upto(101) { |i| eval("s#{i}=s1[0..-1]") }

puts s100.length.to\_s  
puts "Free megabytes " + free\_megs

Output:

Free megabytes 588  
1048576  
Free megabytes 587

Only one meg is used, which is the length of the original string. So,  
by inductive inference, the substrings are only pointers back to the  
original string rather than copies of the data.

Regards,  
Jordan
