# String iterate through regex matches with possition

**URL:** <https://rubytalk.org/t/string-iterate-through-regex-matches-with-possition/66264>\
**Category:** ruby-talk\
**Created:** [11 September 2012 14:44 UTC](https://rubytalk.org/t/string-iterate-through-regex-matches-with-possition/66264 "2012-09-11T14:44:51Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Vicente\_Bosch](https://avatars.discourse-cdn.com/v4/letter/v/76d3ee/32.png) [@Vicente\_Bosch](https://rubytalk.org/u/Vicente_Bosch)\
**Post date:** [11 September 2012 14:44 UTC](https://rubytalk.org/t/string-iterate-through-regex-matches-with-possition/66264/1 "2012-09-11T14:44:51Z")

</div>

Hi,

First of all sorry if this a duplicate question ( I have scanned through  
the last answers regarding regex and didn't get any ideas ).

I am scanning a string in order to detect correctly formed "records" in it.  
A correct record is a "SP" mark followed by "NL" marks (0 or more ) and an  
ending "EP" mark.  
If we find an two EPs without a SP in the middle, two SPs without a EP in  
the middle, or a mark other than "NL" in between the SP and EP marks the  
record is invalid.

"BS HD SP SP EP SP NL EP EP FT BS"

We have the following records:

- SP EP

- SP NL EP

I scan through them and I am able to retrieve them with:

string.scan(/(SP)\s((?:NL\s)\*)(EP)/)

But I am not getting the start and end position of the match inside the  
string ( which I need to retrive data from another place).

Is there any way to scan the string for matches where I get the index  
possition ?

Maybe I should not even be using scan ?

Thanks for your help and time.

Regards,  
V.

---

<div class="post-metadata">

**Author:** ![7stud2](https://avatars.discourse-cdn.com/v4/letter/7/9de053/32.png) [@7stud2](https://rubytalk.org/u/7stud2)\
**Post date:** [11 September 2012 15:25 UTC](https://rubytalk.org/t/string-iterate-through-regex-matches-with-possition/66264/2 "2012-09-11T15:25:32Z")

</div>

Hi,

The MatchData object in $~ has a "offset" method to retrieve the start  
end end offset a capture group. However, I don't understand why you  
capture "SP" and "EP".

string.scan /SP\s((?:NL\s)\*)EP/ do  
&nbsp;&nbsp;p $~.offset 1  
end

> **···**
>
> --  
> Posted via [http://www.ruby-forum.com/](http://www.ruby-forum.com/).

---

<div class="post-metadata">

**Author:** ![Robert\_K1](https://yyz1.discourse-cdn.com/flex029/user_avatar/rubytalk.org/robert_k1/32/1830_2.png) [@Robert\_K1](https://rubytalk.org/u/Robert_K1)\
**Post date:** [11 September 2012 18:26 UTC](https://rubytalk.org/t/string-iterate-through-regex-matches-with-possition/66264/3 "2012-09-11T18:26:45Z")

</div>

See Jan's reply for obtaining the position.

> Maybe I should not even be using scan ?

If you need to process the content in between then you could also use #split:

m = string.split /((?:SP)\s(?:NL\s)\*EP)/

(When #split is used with capturing groups those are retained in the  
resulting array.)

Kind regards

robert

> **···**
>
> On Tue, Sep 11, 2012 at 4:44 PM, Vicente Bosch \<vbosch@gmail.com\> wrote:
> 
> --  
> remember.guy do |as, often| as.you\_can - without end  
> [http://blog.rubybestpractices.com/](http://blog.rubybestpractices.com/)

---

<div class="post-metadata">

**Author:** ![7stud2](https://avatars.discourse-cdn.com/v4/letter/7/9de053/32.png) [@7stud2](https://rubytalk.org/u/7stud2)\
**Post date:** [11 September 2012 20:48 UTC](https://rubytalk.org/t/string-iterate-through-regex-matches-with-possition/66264/4 "2012-09-11T20:48:10Z")

</div>

Vicente Bosch wrote in post #1075485:

> Is there any way to scan the string for matches where I get the index  
> possition ?

str = 'BS HD SP SP EP SP NL EP EP FT BS'

str.scan(/  
&nbsp;&nbsp;&nbsp;&nbsp;SP  
&nbsp;&nbsp;&nbsp;&nbsp;\s\*  
&nbsp;&nbsp;&nbsp;&nbsp;(?:NL)\*  
&nbsp;&nbsp;&nbsp;&nbsp;\s\*  
&nbsp;&nbsp;&nbsp;&nbsp;EP  
/xms) do |match|  
&nbsp;&nbsp;md = Regexp.last\_match  
&nbsp;&nbsp;puts "#{match.inspect} =\> #{md.offset(0)}"  
end

--output:--  
"SP EP" =\> [9, 14]  
"SP NL EP" =\> [15, 23]

> **···**
>
> --  
> Posted via [http://www.ruby-forum.com/\](http://www.ruby-forum.com/%5C).

---

<div class="post-metadata">

**Author:** ![Vicente\_Bosch](https://avatars.discourse-cdn.com/v4/letter/v/76d3ee/32.png) [@Vicente\_Bosch](https://rubytalk.org/u/Vicente_Bosch)\
**Post date:** [12 September 2012 07:55 UTC](https://rubytalk.org/t/string-iterate-through-regex-matches-with-possition/66264/5 "2012-09-12T07:55:41Z")

</div>

Thanks for the answers!! Going to go with Regexp.last\_match 🙂

> **···**
>
> On 11 September 2012 22:48, 7stud -- \<lists@ruby-forum.com\> wrote:
> 
> > Vicente Bosch wrote in post #1075485:  
> > \>  
> > \> Is there any way to scan the string for matches where I get the index  
> > \> possition ?  
> > \>
> > 
> > str = 'BS HD SP SP EP SP NL EP EP FT BS'
> > 
> > str.scan(/  
> > &nbsp;&nbsp;&nbsp;&nbsp;SP  
> > &nbsp;&nbsp;&nbsp;&nbsp;\s\*  
> > &nbsp;&nbsp;&nbsp;&nbsp;(?:NL)\*  
> > &nbsp;&nbsp;&nbsp;&nbsp;\s\*  
> > &nbsp;&nbsp;&nbsp;&nbsp;EP  
> > /xms) do |match|  
> > &nbsp;&nbsp;md = Regexp.last\_match  
> > &nbsp;&nbsp;puts "#{match.inspect} =\> #{md.offset(0)}"  
> > end
> > 
> > --output:--  
> > "SP EP" =\> [9, 14]  
> > "SP NL EP" =\> [15, 23]
> > 
> > --  
> > Posted via [http://www.ruby-forum.com/\](http://www.ruby-forum.com/%5C).
