# General Nokogiri problem

**URL:** <https://rubytalk.org/t/general-nokogiri-problem/53271>\
**Category:** ruby-talk\
**Created:** [7 May 2009 06:45 UTC](https://rubytalk.org/t/general-nokogiri-problem/53271 "2009-05-07T06:45:28Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Srijayanth\_Sridhar](https://avatars.discourse-cdn.com/v4/letter/s/b9bd4f/32.png) [@Srijayanth\_Sridhar](https://rubytalk.org/u/Srijayanth_Sridhar)\
**Post date:** [7 May 2009 06:45 UTC](https://rubytalk.org/t/general-nokogiri-problem/53271/1 "2009-05-07T06:45:28Z")

</div>

Hello,

On several sites(probably malformed HTML/JavaScript/XML/general parsing  
hell) I have the following problem.

For ex:

moonwolf@trantor:~/ruby$ irb  
irb(main):001:0\> ['rubygems','nokogiri','hpricot','open-uri'].each { |r|  
require r }  
=\> ["rubygems", "nokogiri", "hpricot", "open-uri"]  
irb(main):002:0\> doc=Nokogiri(open("[http://maps.google.com/](http://maps.google.com/)"))  
=\> \<?xml version="1.0"?\>  
\<!DOCTYPE html\>  
\<html/\>

irb(main):003:0\> doc/"a"  
=\>

Same with Nokogiri.Hpricot:

irb(main):004:0\> doc=Nokogiri.Hpricot(open("[http://maps.google.com/](http://maps.google.com/)"))  
=\> \<?xml version="1.0"?\>  
\<!DOCTYPE html\>  
\<html/\>

However with regular Hpricot:

irb(main):009:0\> (Hpricot(open("[http://maps.google.com/"))/"a](http://maps.google.com/%22))/%22a)").size  
=\> 53  
(the full post of course is too long, so just showed something simpler)

Hpricot by itself of course works. I tried looking and there's not much by  
way of documentation or blogs on something like this.

Any suggestions/explanations will be welcome as I like Nokogiri's speed very  
much.

I am using:

moonwolf@trantor:~/ruby$ gem list --local | grep -i nokogiri  
nokogiri (1.2.3)  
moonwolf@trantor:~/ruby$ ruby --version  
ruby 1.8.6 (2008-03-03 patchlevel 114) [i686-linux]

Jayanth

---

<div class="post-metadata">

**Author:** ![Aaron\_Patterson1](https://avatars.discourse-cdn.com/v4/letter/a/e0b2c6/32.png) [@Aaron\_Patterson1](https://rubytalk.org/u/Aaron_Patterson1)\
**Post date:** [7 May 2009 07:02 UTC](https://rubytalk.org/t/general-nokogiri-problem/53271/2 "2009-05-07T07:02:40Z")

</div>

> Hello,
> 
> On several sites(probably malformed HTML/JavaScript/XML/general parsing  
> hell) I have the following problem.
> 
> For ex:
> 
> moonwolf@trantor:~/ruby$ irb  
> irb(main):001:0\> ['rubygems','nokogiri','hpricot','open-uri'].each { |r|  
> require r }  
> =\> ["rubygems", "nokogiri", "hpricot", "open-uri"]  
> irb(main):002:0\> doc=Nokogiri(open("[http://maps.google.com/&quot;\](http://maps.google.com/&quot;%5C)))  
> =\> \<?xml version="1.0"?\>  
> \<!DOCTYPE html\>  
> \<html/\>
> 
> irb(main):003:0\> doc/"a"  
> =\>
> 
> Same with Nokogiri.Hpricot:
> 
> irb(main):004:0\> doc=Nokogiri.Hpricot(open("[http://maps.google.com/&quot;\](http://maps.google.com/&quot;%5C)))  
> =\> \<?xml version="1.0"?\>  
> \<!DOCTYPE html\>  
> \<html/\>
> 
> However with regular Hpricot:
> 
> irb(main):009:0\> (Hpricot(open("[http://maps.google.com/&quot;\](http://maps.google.com/&quot;%5C)))/"a").size  
> =\> 53  
> (the full post of course is too long, so just showed something simpler)
> 
> Hpricot by itself of course works. I tried looking and there's not much by  
> way of documentation or blogs on something like this.
> 
> Any suggestions/explanations will be welcome as I like Nokogiri's speed very  
> much.

Nokogiri detects the XML header and parses it as XML. If you force it  
to use the HTML parser, you may be more successfull:

&nbsp;&nbsp;\>\> (Nokogiri::HTML(open("[http://maps.google.com/&quot;\](http://maps.google.com/&quot;%5C)))/'a').length  
&nbsp;&nbsp;=\> 53

> **···**
>
> On Thu, May 07, 2009 at 03:45:28PM +0900, Srijayanth Sridhar wrote:  
> &nbsp;&nbsp;\>\>
> 
> --  
> Aaron Patterson  
> [http://tenderlovemaking.com/](http://tenderlovemaking.com/)

---

<div class="post-metadata">

**Author:** ![Srijayanth\_Sridhar](https://avatars.discourse-cdn.com/v4/letter/s/b9bd4f/32.png) [@Srijayanth\_Sridhar](https://rubytalk.org/u/Srijayanth_Sridhar)\
**Post date:** [7 May 2009 07:05 UTC](https://rubytalk.org/t/general-nokogiri-problem/53271/3 "2009-05-07T07:05:33Z")

</div>

Thanks Aaron.

Jayanth

> **···**
>
> On Thu, May 7, 2009 at 12:32 PM, Aaron Patterson \<aaron@tenderlovemaking.com \> wrote:
> 
> > On Thu, May 07, 2009 at 03:45:28PM +0900, Srijayanth Sridhar wrote:  
> > \> Hello,  
> > \>  
> > \> On several sites(probably malformed HTML/JavaScript/XML/general parsing  
> > \> hell) I have the following problem.  
> > \>  
> > \> For ex:  
> > \>  
> > \> moonwolf@trantor:~/ruby$ irb  
> > \> irb(main):001:0\> ['rubygems','nokogiri','hpricot','open-uri'].each { |r|  
> > \> require r }  
> > \> =\> ["rubygems", "nokogiri", "hpricot", "open-uri"]  
> > \> irb(main):002:0\> doc=Nokogiri(open("[http://maps.google.com/&quot;\](http://maps.google.com/&quot;%5C)))  
> > \> =\> \<?xml version="1.0"?\>  
> > \> \<!DOCTYPE html\>  
> > \> \<html/\>  
> > \>  
> > \> irb(main):003:0\> doc/"a"  
> > \> =\>  
> > \>  
> > \> Same with Nokogiri.Hpricot:  
> > \>  
> > \> irb(main):004:0\> doc=Nokogiri.Hpricot(open("[http://maps.google.com/&quot;\](http://maps.google.com/&quot;%5C)))  
> > \> =\> \<?xml version="1.0"?\>  
> > \> \<!DOCTYPE html\>  
> > \> \<html/\>  
> > \>  
> > \> However with regular Hpricot:  
> > \>  
> > \> irb(main):009:0\> (Hpricot(open("[http://maps.google.com/&quot;\](http://maps.google.com/&quot;%5C)))/"a").size  
> > \> =\> 53  
> > \> (the full post of course is too long, so just showed something simpler)  
> > \>  
> > \>  
> > \> Hpricot by itself of course works. I tried looking and there's not much  
> > by  
> > \> way of documentation or blogs on something like this.  
> > \>  
> > \> Any suggestions/explanations will be welcome as I like Nokogiri's speed  
> > very  
> > \> much.
> > 
> > Nokogiri detects the XML header and parses it as XML. If you force it  
> > to use the HTML parser, you may be more successfull:
> > 
> > \>\> (Nokogiri::HTML(open("[http://maps.google.com/&quot;\](http://maps.google.com/&quot;%5C)))/'a').length  
> > =\> 53  
> > \>\>
> > 
> > --  
> > Aaron Patterson  
> > [http://tenderlovemaking.com/](http://tenderlovemaking.com/)

---

<div class="post-metadata">

**Author:** ![Srijayanth\_Sridhar](https://avatars.discourse-cdn.com/v4/letter/s/b9bd4f/32.png) [@Srijayanth\_Sridhar](https://rubytalk.org/u/Srijayanth_Sridhar)\
**Post date:** [7 May 2009 07:08 UTC](https://rubytalk.org/t/general-nokogiri-problem/53271/4 "2009-05-07T07:08:14Z")

</div>

Whoops,

irb(main):015:0\> (Nokogiri::HTML(open("[http://maps.google.com/](http://maps.google.com/)  
"))/'a').length  
=\> 0

Not sure what the deal is.

Jayanth

> **···**
>
> On Thu, May 7, 2009 at 12:35 PM, Srijayanth Sridhar \<srijayanth@gmail.com\>wrote:
> 
> > Thanks Aaron.
> > 
> > Jayanth
> > 
> > On Thu, May 7, 2009 at 12:32 PM, Aaron Patterson \< \> aaron@tenderlovemaking.com\> wrote:
> > 
> > > On Thu, May 07, 2009 at 03:45:28PM +0900, Srijayanth Sridhar wrote:  
> > > \> Hello,  
> > > \>  
> > > \> On several sites(probably malformed HTML/JavaScript/XML/general parsing  
> > > \> hell) I have the following problem.  
> > > \>  
> > > \> For ex:  
> > > \>  
> > > \> moonwolf@trantor:~/ruby$ irb  
> > > \> irb(main):001:0\> ['rubygems','nokogiri','hpricot','open-uri'].each { |r|  
> > > \> require r }  
> > > \> =\> ["rubygems", "nokogiri", "hpricot", "open-uri"]  
> > > \> irb(main):002:0\> doc=Nokogiri(open("[http://maps.google.com/&quot;\](http://maps.google.com/&quot;%5C)))  
> > > \> =\> \<?xml version="1.0"?\>  
> > > \> \<!DOCTYPE html\>  
> > > \> \<html/\>  
> > > \>  
> > > \> irb(main):003:0\> doc/"a"  
> > > \> =\>  
> > > \>  
> > > \> Same with Nokogiri.Hpricot:  
> > > \>  
> > > \> irb(main):004:0\> doc=Nokogiri.Hpricot(open("[http://maps.google.com/&quot;\](http://maps.google.com/&quot;%5C)))  
> > > \> =\> \<?xml version="1.0"?\>  
> > > \> \<!DOCTYPE html\>  
> > > \> \<html/\>  
> > > \>  
> > > \> However with regular Hpricot:  
> > > \>  
> > > \> irb(main):009:0\> (Hpricot(open("[http://maps.google.com/&quot;\](http://maps.google.com/&quot;%5C)))/"a").size  
> > > \> =\> 53  
> > > \> (the full post of course is too long, so just showed something simpler)  
> > > \>  
> > > \>  
> > > \> Hpricot by itself of course works. I tried looking and there's not much  
> > > by  
> > > \> way of documentation or blogs on something like this.  
> > > \>  
> > > \> Any suggestions/explanations will be welcome as I like Nokogiri's speed  
> > > very  
> > > \> much.
> > > 
> > > Nokogiri detects the XML header and parses it as XML. If you force it  
> > > to use the HTML parser, you may be more successfull:
> > > 
> > > \>\> (Nokogiri::HTML(open("[http://maps.google.com/&quot;\](http://maps.google.com/&quot;%5C)))/'a').length  
> > > =\> 53  
> > > \>\>
> > > 
> > > --  
> > > Aaron Patterson  
> > > [http://tenderlovemaking.com/](http://tenderlovemaking.com/)
