# scRUBYt! 0.3.1 released

**URL:** <https://rubytalk.org/t/scrubyt-0-3-1-released/37939>\
**Category:** ruby-talk\
**Created:** [29 May 2007 19:34 UTC](https://rubytalk.org/t/scrubyt-0-3-1-released/37939 "2007-05-29T19:34:19Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Peter\_Szinek](https://avatars.discourse-cdn.com/v4/letter/p/5fc32e/32.png) [@Peter\_Szinek](https://rubytalk.org/u/Peter_Szinek)\
**Post date:** [29 May 2007 19:34 UTC](https://rubytalk.org/t/scrubyt-0-3-1-released/37939/1 "2007-05-29T19:34:19Z")

</div>

Hello all,

scRUBYt! version 0.3.1 has been released with a plenty of new features  
and bugfixes based on your feedback. Enjoy!

> **···**
>
> # ============ What’s this?
> 
> scRUBYt! is a very easy to learn and use, yet powerful Web scraping  
> framework based on Hpricot and mechanize. It’s purpose is to free you  
> from the drudgery of web page crawling, looking up HTML tags,  
> attributes, XPaths, form names and other typical low-level web scraping  
> woes by figuring these out from your examples copy’n’pasted from the Web  
> page.
> 
> # =========== What’s new?
> 
> [NEW] complete rewrite of the output system, creating  
> a solid foundation for more robust output functions  
> (credit: Neelance)  
> [NEW] logging - no annoying puts messages anymore!  
> (credit: Tim Fletcher)  
> [NEW] can index an example - e.g.  
> link ‘more[5]’  
> semantics: give me the 6th element with the text ‘link’  
> [NEW] can use XPath checking an attribute value, like  
> “//div[@id=‘content’]”  
> [NEW] default values for missing elements (first version was done in  
> 0.2.8 but it did not work for all cases)  
> [NEW] possibility to click button with it’s text (instead of it’s index)  
> (credit: Nick Merwin)  
> [NEW] clicking radio buttons  
> [NEW] can click on image buttons (by specifying the name of the button)  
> [NEW] possibility to extract an URL with one step, like so:  
> link ‘The Difference/@href’  
> i.e. give me the href attribute of the element matched by the  
> example ‘The Difference’  
> [NEW] new way to match an element of the page:  
> div ‘div[The Difference]’  
> means ‘return the div which contains the string “The Difference”’.  
> This is useful if the XPath of the element is non-constant across  
> the same site (e.g.sometimes a banner or add is added, sometimes  
> not etc.)  
> [NEW] Clicking image maps; At the moment this is achieved by specifying  
> an index, like  
> click\_image\_map 3  
> which means click the 4th link in the image map  
> [FIX] Replacing \240 (&nbsp;) with space in the preprocessing phase  
> automatically  
> [FIX] Fixed: correctly downloading image if the src  
> attribute had a leading space, as in  
> ![](https://yyz1.discourse-cdn.com/flex029/files/downloads/images/image.jpg)  
> [FIX] Other misc fixes - a ton of them!
> 
> # ======== Comments
> 
> The win32 version is just being built as I am writing this, so it will  
> be available soon.
> 
> Please keep the feedback coming - bug reports, questions, suggestions  
> are warmly welcome at the scRUBYt! forum - [http://agora.scrubyt.org](http://agora.scrubyt.org).
> 
> Cheers,  
> The scRUBYt! team - [http://scrubyt.org](http://scrubyt.org)
> 
> –~--~---------~–~----~------------~-------~–~----~  
> You received this message because you are subscribed to the Google Groups “Ruby on Rails: Talk” group.  
> To post to this group, send email to rubyonrails-talk-/JYPxA39Uh5TLH3MbocFFw@public.gmane.org  
> To unsubscribe from this group, send email to rubyonrails-talk-unsubscribe-/JYPxA39Uh5TLH3MbocFFw@public.gmane.org  
> For more options, visit this group at [http://groups.google.com/group/rubyonrails-talk?hl=en](http://groups.google.com/group/rubyonrails-talk?hl=en)  
> -~----------~----~----~----~------~----~------~–~—

---

<div class="post-metadata">

**Author:** ![al\_batuul](https://avatars.discourse-cdn.com/v4/letter/a/f08c70/32.png) [@al\_batuul](https://rubytalk.org/u/al_batuul)\
**Post date:** [29 May 2007 19:46 UTC](https://rubytalk.org/t/scrubyt-0-3-1-released/37939/2 "2007-05-29T19:46:16Z")

</div>

Waiting eagerly for the window version  
al\_batuul

Peter Szinek \<[peter@rubyrailways.com](mailto:peter@rubyrailways.com)\> wrote: Hello all,

scRUBYt! version 0.3.1 has been released with a plenty of new features  
and bugfixes based on your feedback. Enjoy!

> **···**
>
> # ============ What's this?
> 
> scRUBYt! is a very easy to learn and use, yet powerful Web scraping  
> framework based on Hpricot and mechanize. It's purpose is to free you  
> from the drudgery of web page crawling, looking up HTML tags,  
> attributes, XPaths, form names and other typical low-level web scraping  
> woes by figuring these out from your examples copy'n'pasted from the Web  
> page.
> 
> # =========== What's new?
> 
> [NEW] complete rewrite of the output system, creating  
> a solid foundation for more robust output functions  
> (credit: Neelance)  
> [NEW] logging - no annoying puts messages anymore!  
> (credit: Tim Fletcher)  
> [NEW] can index an example - e.g.  
> link 'more[5]'  
> semantics: give me the 6th element with the text 'link'  
> [NEW] can use XPath checking an attribute value, like  
> "//div[@id='content']"  
> [NEW] default values for missing elements (first version was done in  
> 0.2.8 but it did not work for all cases)  
> [NEW] possibility to click button with it's text (instead of it's index)  
> (credit: Nick Merwin)  
> [NEW] clicking radio buttons  
> [NEW] can click on image buttons (by specifying the name of the button)  
> [NEW] possibility to extract an URL with one step, like so:  
> link 'The Difference/@href'  
> i.e. give me the href attribute of the element matched by the  
> example 'The Difference'  
> [NEW] new way to match an element of the page:  
> div 'div[The Difference]'  
> means 'return the div which contains the string "The Difference"'.  
> This is useful if the XPath of the element is non-constant across  
> the same site (e.g.sometimes a banner or add is added, sometimes  
> not etc.)  
> [NEW] Clicking image maps; At the moment this is achieved by specifying  
> an index, like  
> click\_image\_map 3  
> which means click the 4th link in the image map  
> [FIX] Replacing \240 ( ) with space in the preprocessing phase  
> automatically  
> [FIX] Fixed: correctly downloading image if the src  
> attribute had a leading space, as in
> 
> [FIX] Other misc fixes - a ton of them!
> 
> # ======== Comments
> 
> The win32 version is just being built as I am writing this, so it will  
> be available soon.
> 
> Please keep the feedback coming - bug reports, questions, suggestions  
> are warmly welcome at the scRUBYt! forum - [http://agora.scrubyt.org](http://agora.scrubyt.org).
> 
> Cheers,  
> The scRUBYt! team - [http://scrubyt.org](http://scrubyt.org)
> 
> ---------------------------------  
> Luggage? GPS? Comic books?  
> Check out fitting gifts for grads at Yahoo! Search.
