# Performance comparison between screen scrapers

**URL:** <https://rubytalk.org/t/performance-comparison-between-screen-scrapers/34375>\
**Category:** ruby-talk\
**Created:** [11 January 2007 08:36 UTC](https://rubytalk.org/t/performance-comparison-between-screen-scrapers/34375 "2007-01-11T08:36:44Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Conrad\_Chu](https://avatars.discourse-cdn.com/v4/letter/c/eada6e/32.png) [@Conrad\_Chu](https://rubytalk.org/u/Conrad_Chu)\
**Post date:** [11 January 2007 08:36 UTC](https://rubytalk.org/t/performance-comparison-between-screen-scrapers/34375/1 "2007-01-11T08:36:44Z")

</div>

Does anyone know how the following screen scrapers perform against one  
another?

\* ScrAPI  
\* RubyfulSoup  
\* HTree  
\* Hpricot

I'm trying to write up a tool where a person enters in a URL, and I use  
an AJAX call to scrape the contents of that URL for title, description,  
etc. So speed is really important (I suppose, regular expressions would  
be the fastest, but I need something that is tree-based and supports  
HTML tidying)

Thanks  
Conrad

> **···**
>
> --  
> Posted via [http://www.ruby-forum.com/](http://www.ruby-forum.com/).

---

<div class="post-metadata">

**Author:** ![Jan\_Svitok](https://avatars.discourse-cdn.com/v4/letter/j/cab0a1/32.png) [@Jan\_Svitok](https://rubytalk.org/u/Jan_Svitok)\
**Post date:** [11 January 2007 09:51 UTC](https://rubytalk.org/t/performance-comparison-between-screen-scrapers/34375/2 "2007-01-11T09:51:16Z")

</div>

There was a comparision done on this list some time ago. Search for lib names.

> **···**
>
> On 1/11/07, Conrad Chu \<conradchu@conradchu.com\> wrote:
> 
> > Does anyone know how the following screen scrapers perform against one  
> > another?
> > 
> > \* ScrAPI  
> > \* RubyfulSoup  
> > \* HTree  
> > \* Hpricot
> > 
> > I'm trying to write up a tool where a person enters in a URL, and I use  
> > an AJAX call to scrape the contents of that URL for title, description,  
> > etc. So speed is really important (I suppose, regular expressions would  
> > be the fastest, but I need something that is tree-based and supports  
> > HTML tidying)
> > 
> > Thanks  
> > Conrad

---

<div class="post-metadata">

**Author:** ![Ross\_Bamford2](https://avatars.discourse-cdn.com/v4/letter/r/e47774/32.png) [@Ross\_Bamford2](https://rubytalk.org/u/Ross_Bamford2)\
**Post date:** [11 January 2007 10:25 UTC](https://rubytalk.org/t/performance-comparison-between-screen-scrapers/34375/3 "2007-01-11T10:25:06Z")

</div>

I don't know about ScrAPI or HTree, but I recently blogged an informal benchmark run between Rubyful Soup, Hpricot, and the (still developmental) libxml2 HTML parser binding in Libxml-ruby. It's at:

&nbsp;&nbsp;[http://cloverhead.blogspot.com/2006/12/bit-of-benchmarking.html](http://cloverhead.blogspot.com/2006/12/bit-of-benchmarking.html)

> **···**
>
> On Thu, 11 Jan 2007 08:36:44 -0000, Conrad Chu \<conradchu@conradchu.com\> wrote:
> 
> > Does anyone know how the following screen scrapers perform against one  
> > another?
> > 
> > \* ScrAPI  
> > \* RubyfulSoup  
> > \* HTree  
> > \* Hpricot
> > 
> > I'm trying to write up a tool where a person enters in a URL, and I use  
> > an AJAX call to scrape the contents of that URL for title, description,  
> > etc. So speed is really important (I suppose, regular expressions would  
> > be the fastest, but I need something that is tree-based and supports  
> > HTML tidying)
> > 
> > Thanks  
> > Conrad
> 
> --  
> Ross Bamford - rosco@roscopeco.remove.co.uk

---

<div class="post-metadata">

**Author:** ![Interfecus](https://avatars.discourse-cdn.com/v4/letter/i/d78d45/32.png) [@Interfecus](https://rubytalk.org/u/Interfecus)\
**Post date:** [13 January 2007 13:15 UTC](https://rubytalk.org/t/performance-comparison-between-screen-scrapers/34375/4 "2007-01-13T13:15:13Z")

</div>

Conrad Chu wrote:

> Does anyone know how the following screen scrapers perform against one  
> another?
> 
> \* ScrAPI  
> \* RubyfulSoup  
> \* HTree  
> \* Hpricot
> 
> I'm trying to write up a tool where a person enters in a URL, and I use  
> an AJAX call to scrape the contents of that URL for title, description,  
> etc. So speed is really important (I suppose, regular expressions would  
> be the fastest, but I need something that is tree-based and supports  
> HTML tidying)
> 
> Thanks  
> Conrad
> 
> --  
> Posted via [http://www.ruby-forum.com/\](http://www.ruby-forum.com/%5C).

I haven't used them all but Hpricot is fast (the parser is written in C  
with Ragel), error tolerant and perfect for this task. Take a look at  
its website for a guide on how to use it.
