# JSON.parse and unicode escape?

**URL:** <https://rubytalk.org/t/json-parse-and-unicode-escape/48637>\
**Category:** ruby-talk\
**Created:** [26 August 2008 23:26 UTC](https://rubytalk.org/t/json-parse-and-unicode-escape/48637 "2008-08-26T23:26:01Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jonathan\_Rochkind](https://avatars.discourse-cdn.com/v4/letter/j/7bcc69/32.png) [@Jonathan\_Rochkind](https://rubytalk.org/u/Jonathan_Rochkind)\
**Post date:** [26 August 2008 23:26 UTC](https://rubytalk.org/t/json-parse-and-unicode-escape/48637/1 "2008-08-26T23:26:01Z")

</div>

The documentation for the ruby JSON classes ([http://json.rubyforge.org/](http://json.rubyforge.org/))  
implies that it handles unicode escaping fine. But I'm having trouble  
with parsing JSON with a unicode escape sequence in it. I am using the  
'ext' parser (JSON::Ext::Parser) not the 'pure' parser. version 1.1.2,  
which appears to still be the latest.

Here is some test JSON, that's actually an excerpt from some JSON  
returned to me by a third party web service. Finally boiled it down to  
the simplest demonstration case. I saved it in a file, but here's what's  
in the text file:

> **···**
>
> # ===== { "key": 'something \x26 more' }
> 
> I believe that is valid json, containing an escaped unicode char? But  
> JSON.parse on that string throws, complaining:
> 
> JSON::ParserError: unexpected token at '{ "summary": ' \u0026 ' }
> 
> I have verified it is the /x26 that's doing it. It doesn't like \x  
> escaped unicode.
> 
> Am I doing something wrong? Is the JSON I am receiving from the third  
> party bad somehow? This is such a widely used library that I'd be  
> surprised if it's broken and can't parse input including unicode escape  
> sequences... but that's what it looks like to me. Feedback?  
> --  
> Posted via [http://www.ruby-forum.com/](http://www.ruby-forum.com/).

---

<div class="post-metadata">

**Author:** ![pwever](https://avatars.discourse-cdn.com/v4/letter/p/87869e/32.png) [@pwever](https://rubytalk.org/u/pwever)\
**Post date:** [1 October 2008 04:18 UTC](https://rubytalk.org/t/json-parse-and-unicode-escape/48637/2 "2008-10-01T04:18:59Z")

</div>

I am running into what seems to be a related problem with the  
following code:

irb

> > require 'json'

=\> true

> > JSON.parse('{"s":"\uddb0"}')

JSON::ParserError: source sequence is illegal/malformed near uddb0"}  
&nbsp;&nbsp;from /Library/Ruby/Gems/1.8/gems/json-1.1.3/lib/json/common.rb:122:in  
`parse'  
&nbsp;&nbsp;from /Library/Ruby/Gems/1.8/gems/json-1.1.3/lib/json/common.rb:122:in  
`parse'  
&nbsp;&nbsp;from (irb):2

> >

I don't know enough about unicode to really understand what is being  
escaped here, but the following unicode characters, very close in  
range (I assume) do not throw an error:  
"\ucdb0", "\uedb0", "\ud7b0"

I also validated the JSON string ('{"s":"\uddb0"}') successfully at  
[http://www.jsonlint.com/](http://www.jsonlint.com/) and in Python.

Any ideas of what might be the problem?  
Are there any alternative JSON parsers for ruby?

Thank you very much // pascal

> **···**
>
> from :0

---

<div class="post-metadata">

**Author:** ![Rob\_Biedenharn1](https://avatars.discourse-cdn.com/v4/letter/r/57b2e6/32.png) [@Rob\_Biedenharn1](https://rubytalk.org/u/Rob_Biedenharn1)\
**Post date:** [1 October 2008 04:29 UTC](https://rubytalk.org/t/json-parse-and-unicode-escape/48637/3 "2008-10-01T04:29:10Z")

</div>

That's not valid Unicode. See:

> **[UDC00.pdf](https://www.unicode.org/charts/PDF/UDC00.pdf)**
>
> 48.08 KB

You can only have that code point in UTF-16

-Rob

> **···**
>
> On Oct 1, 2008, at 12:18 AM, pwever wrote:
> 
> > I am running into what seems to be a related problem with the  
> > following code:
> > 
> > irb
> > 
> > > > require 'json'
> > 
> > =\> true
> > 
> > > > JSON.parse('{"s":"\uddb0"}')
> > 
> > JSON::ParserError: source sequence is illegal/malformed near uddb0"}  
> > &nbsp;&nbsp;from /Library/Ruby/Gems/1.8/gems/json-1.1.3/lib/json/common.rb:122:in  
> > `parse'  
> > &nbsp;&nbsp;from /Library/Ruby/Gems/1.8/gems/json-1.1.3/lib/json/common.rb:122:in  
> > `parse'  
> > &nbsp;&nbsp;from (irb):2  
> > &nbsp;&nbsp;from :0
> > 
> > > >
> > 
> > I don't know enough about unicode to really understand what is being  
> > escaped here, but the following unicode characters, very close in  
> > range (I assume) do not throw an error:  
> > "\ucdb0", "\uedb0", "\ud7b0"
> > 
> > I also validated the JSON string ('{"s":"\uddb0"}') successfully at  
> > [http://www.jsonlint.com/](http://www.jsonlint.com/) and in Python.
> > 
> > Any ideas of what might be the problem?  
> > Are there any alternative JSON parsers for ruby?
> > 
> > Thank you very much // pascal

---

<div class="post-metadata">

**Author:** ![Pascal\_Wever](https://avatars.discourse-cdn.com/v4/letter/p/47e85d/32.png) [@Pascal\_Wever](https://rubytalk.org/u/Pascal_Wever)\
**Post date:** [1 October 2008 17:55 UTC](https://rubytalk.org/t/json-parse-and-unicode-escape/48637/4 "2008-10-01T17:55:40Z")

</div>

That makes a lot of sense. Thanks for the clarification regarding the  
unicode range.

Since I don't have control over the JSON source, I would like to try to  
parse the JSON even if it results in a malformed unicode string. So  
today I tried switching from 'json' to the 'ruby-json' library. After  
some searching online, I didn't find any documentation on how to use it  
though. Primarily I don't know how to include or require it.

require 'ruby-json'  
require 'rubyjson'

don't seem to work?  
Any ideas are appreaciated.  
Thank you very much  
// pascal

> **···**
>
> --  
> Posted via [http://www.ruby-forum.com/](http://www.ruby-forum.com/).
