# Strings vs arrays

**URL:** <https://rubytalk.org/t/strings-vs-arrays/19440>\
**Category:** ruby-talk\
**Created:** [9 July 2005 12:46 UTC](https://rubytalk.org/t/strings-vs-arrays/19440 "2005-07-09T12:46:56Z")\
**Posts on this page:** 5\
**Page:** 2

<div class="post-metadata">

**Author:** ![David\_A\_Black3](https://avatars.discourse-cdn.com/v4/letter/d/6a8cbe/32.png) [@David\_A\_Black3](https://rubytalk.org/u/David_A_Black3)\
**Post date:** [9 July 2005 21:01 UTC](https://rubytalk.org/t/strings-vs-arrays/19440/21 "2005-07-09T21:01:04Z")

</div>

Hi --

> Daniel Brockman wrote:
> 
> > Whatever String# and all the other String methods index, of course.
> 
> Depending on the parameter you pass, # can return a String or an Integer.
> 
> > > There is no clear notion of an "Element" in a String.
> > 
> > If this is true, then we have a serious problem. Before doing much  
> > anything about its API, we need to decide whether String is a byte  
> > array or a character array. (Presumably, matz & co. already have.)

My understanding (sneaking in a reply to Daniel's post in this reply  
🙂 was that the conceptual and design decision was that Strings are  
not arrays, and are therefore not obliged or constrained to have an  
Array-like API (any more than arrays are obliged to have a String-like  
API).

> There's no such thing as a character in Ruby. (See any discussion on Unicode, etc. in Ruby.) Strings are Objects (stored in C as char\*s, I'd guess). Call the right methods on them, and you can get an Integer representing the byte value at a given position ("Hello"[0]), or another String object representing some manipulation of the String ("Hello"[0..0]). Those are your only means of inspection. That was just a long-winded way of saying "this is true."
> 
> Anything I didn't reply to, I probably agree with. Since the String methods don't have a consistent notion of an "element," it doesn't seem it would hurt to choose whichever notion we want for a potential #shift method.

That's true only if there's an imperative to have a String instance  
method called "shift". I don't think there is. Maybe a left  
chop/chomp operation would be a good idea, but I think it should be  
called lchop, which would be consistent with other string method  
naming (rather than with array method naming).

David

> **···**
>
> On Sun, 10 Jul 2005, Devin Mullins wrote:
> 
> --  
> David A. Black  
> dblack@wobblini.net

---

<div class="post-metadata">

**Author:** ![Daniel\_Brockman](https://avatars.discourse-cdn.com/v4/letter/d/ed8c4c/32.png) [@Daniel\_Brockman](https://rubytalk.org/u/Daniel_Brockman)\
**Post date:** [9 July 2005 22:34 UTC](https://rubytalk.org/t/strings-vs-arrays/19440/22 "2005-07-09T22:34:39Z")

</div>

Wow, some mixup --- Daniel, David, Devin and Levin. 🙂

&nbsp;&nbsp;\> There is no clear notion of an "Element" in a String.

&nbsp;&nbsp;\> If this is true, then we have a serious problem. Before  
&nbsp;&nbsp;\> doing much anything about its API, we need to decide whether  
&nbsp;&nbsp;\> String is a byte array or a character array. (Presumably,  
&nbsp;&nbsp;\> matz & co. already have.)

&nbsp;&nbsp;\> My understanding was that the conceptual and design decision  
&nbsp;&nbsp;\> was that Strings are not arrays, and are therefore not  
&nbsp;&nbsp;\> obliged or constrained to have an Array-like API (any more  
&nbsp;&nbsp;\> than arrays are obliged to have a String-like API).

To me, the term \_array\_ means ``sequence of elements stored in a  
contiguous chunk of memory,'' just like the term \_list\_ means  
``sequence of elements stored in a linked list of cells.''

As long as strings are implemented as arrays, I reserve the right to  
refer to them as such, regardless of whether String \< Array.  
(Of course, I won't do it needlessly since it is confusing.)

In Haskell, strings are not arrays, but lists. In many languages,  
strings are immutable, which deemphasises their nature as arrays.  
In Ruby, however, strings are mutable arrays (of bytes?). So far I  
have not been able to understand the desire to obfuscate this fact.

What's the harm of admitting that strings are arrays? What's the harm  
of making them at least \_quack\_ alike?

&nbsp;&nbsp;\> Since the String methods don't have a consistent notion of an  
&nbsp;&nbsp;\> "element," it doesn't seem it would hurt to choose whichever  
&nbsp;&nbsp;\> notion we want for a potential #shift method.

&nbsp;&nbsp;\> That's true only if there's an imperative to have a String  
&nbsp;&nbsp;\> instance method called "shift". I don't think there is.

Why not? This thread started with an example of the need for one.  
Of course you can use string.slice!(0), but the same goes for arrays.

The terms `shift' and `unshift' are general and well-understood.  
I can't see any reason why they shouldn't be applied to strings.

&nbsp;&nbsp;\> Maybe a left chop/chomp operation would be a good idea, but I  
&nbsp;&nbsp;\> think it should be called lchop, which would be consistent  
&nbsp;&nbsp;\> with other string method naming (rather than with array  
&nbsp;&nbsp;\> method naming).

I fail to see the point in that. String#chop is meant to chop off  
end-of-line characters; String#lchop wouldn't be. So using that name  
would be \_inconsistent\_ with other string method naming.

> **···**
>
> --  
> Daniel Brockman \<[daniel@brockman.se](mailto:daniel@brockman.se)\>
> 
> &nbsp;&nbsp;&nbsp;&nbsp;So really, we all have to ask ourselves:  
> &nbsp;&nbsp;&nbsp;&nbsp;Am I waiting for RMS to do this? --TTN.

---

<div class="post-metadata">

**Author:** ![David\_A\_Black3](https://avatars.discourse-cdn.com/v4/letter/d/6a8cbe/32.png) [@David\_A\_Black3](https://rubytalk.org/u/David_A_Black3)\
**Post date:** [10 July 2005 13:27 UTC](https://rubytalk.org/t/strings-vs-arrays/19440/23 "2005-07-10T13:27:54Z")

</div>

Hi --

> \> Maybe a left chop/chomp operation would be a good idea, but I  
> \> think it should be called lchop, which would be consistent  
> \> with other string method naming (rather than with array  
> \> method naming).
> 
> I fail to see the point in that. String#chop is meant to chop off  
> end-of-line characters; String#lchop wouldn't be. So using that name  
> would be \_inconsistent\_ with other string method naming.

String#chop chops off the rightmost character:

&nbsp;&nbsp;&nbsp;irb(main):001:0\> "abc".chop  
&nbsp;&nbsp;&nbsp;=\> "ab"

You may be thinking of "chomp", which is a specialized "chop"  
operating only on newline characters.

So the idea of lchop would be to serve as a left-hand equivalent of  
chop.

David

> **···**
>
> On Sun, 10 Jul 2005, Daniel Brockman wrote:
> 
> --  
> David A. Black  
> dblack@wobblini.net

---

<div class="post-metadata">

**Author:** ![Daniel\_Brockman](https://avatars.discourse-cdn.com/v4/letter/d/ed8c4c/32.png) [@Daniel\_Brockman](https://rubytalk.org/u/Daniel_Brockman)\
**Post date:** [10 July 2005 16:08 UTC](https://rubytalk.org/t/strings-vs-arrays/19440/24 "2005-07-10T16:08:37Z")

</div>

"David A. Black" \<dblack@wobblini.net\> writes:

> String#chop chops off the rightmost character:
> 
> &nbsp;&nbsp;&nbsp;irb(main):001:0\> "abc".chop  
> &nbsp;&nbsp;&nbsp;=\> "ab"

Except if the string ends with a CRLF pair:

&nbsp;&nbsp;&nbsp;"abc\r\n".chop #=\> "abc"

> You may be thinking of "chomp", which is a specialized "chop"  
> operating only on newline characters.

If you read the docstrings, you get the impression that String#chop  
is more-or-less deprecated in favor of the ``safer'' String#chomp:

&nbsp;&nbsp;&nbsp;+String#chomp+ is ofter a safer alternative, as it leaves  
&nbsp;&nbsp;&nbsp;the string unchanged if it doesn't end in a record separator.

> So the idea of lchop would be to serve as a left-hand equivalent  
> of chop.

So I suppose if the string starts with a CRLF pair, String#lchop would  
chop off two characters from the left?

Why not go all the way and let all string methods treat CRLF pairs as  
single characters?

I think it's a problem that strings are the only way to go for raw  
byte arrays in Ruby, yet

&nbsp;&nbsp;\* strings lack a few random useful array methods

&nbsp;&nbsp;\* the string methods are not binary safe.

> **···**
>
> --  
> Daniel Brockman \<daniel@brockman.se\>

---

<div class="post-metadata">

**Author:** ![David\_A\_Black3](https://avatars.discourse-cdn.com/v4/letter/d/6a8cbe/32.png) [@David\_A\_Black3](https://rubytalk.org/u/David_A_Black3)\
**Post date:** [10 July 2005 20:21 UTC](https://rubytalk.org/t/strings-vs-arrays/19440/25 "2005-07-10T20:21:42Z")

</div>

Hi --

> "David A. Black" \<dblack@wobblini.net\> writes:
> 
> > String#chop chops off the rightmost character:
> > 
> > &nbsp;&nbsp;&nbsp;irb(main):001:0\> "abc".chop  
> > &nbsp;&nbsp;&nbsp;=\> "ab"
> 
> Except if the string ends with a CRLF pair:
> 
> &nbsp;&nbsp;"abc\r\n".chop #=\> "abc"
> 
> > You may be thinking of "chomp", which is a specialized "chop"  
> > operating only on newline characters.
> 
> If you read the docstrings, you get the impression that String#chop  
> is more-or-less deprecated in favor of the ``safer'' String#chomp:
> 
> &nbsp;&nbsp;+String#chomp+ is ofter a safer alternative, as it leaves  
> &nbsp;&nbsp;the string unchanged if it doesn't end in a record separator.

I believe that means safer in the sense that if you're going through,  
say, lines in a file, and for some reason there's no \n at the end of  
the last line, you won't accidentally cut off a non-\n character.

In the general case, #chop can't be deprecated in favor of #chomp,  
because #chomp doesn't offer the same functionality (chopping off the  
last character).

> > So the idea of lchop would be to serve as a left-hand equivalent  
> > of chop.
> 
> So I suppose if the string starts with a CRLF pair, String#lchop would  
> chop off two characters from the left?

That's a good question. One could argue that the only reason they are  
treated together in the first place is that they represent the more  
abstract concept "newline" -- and that where they aren't representing  
that concept, they should be treated separately. Or one could go for  
the complete symmetry approach. I guess I'd tend to favor the former  
notion, since the idea of left-end/right-end is already irreducibly  
asymmetrical in a left-to-right writing system. (Though then there's  
the matter of what would happen given a right-to-left writing system,  
etc.)

> Why not go all the way and let all string methods treat CRLF pairs as  
> single characters?

See above -- there's no magic association between those two  
characters, just the historical fact of their serving the newline  
role, and the practical need to acknowledge that role. I don't think  
there would be any advantage to, say, having String#count combine  
them, etc. (though of course there would have been an advantage to  
global agreement several decades ago on how to represent newline on  
various platforms 🙂

> I think it's a problem that strings are the only way to go for raw  
> byte arrays in Ruby, yet
> 
> \* strings lack a few random useful array methods
> 
> \* the string methods are not binary safe.

String, like Hash, raises interesting questions about the relation  
between itself, Array, and Enumerable. It's interesting that  
String#to\_a breaks the string into lines as opposed to characters or  
bytes. That's certainly a behavior one would not expect if the  
"arrayness" of strings resided strictly in their status as ordered  
collections of bytes. On the other hand, they are ordered collections  
of characters of bytes 🙂 I still find myself expecting String#each  
to go bytewise. But the fact that these different objects don't map  
exactly on to each other is, I think, one of the points of having a  
higher and separate abstraction like Enumerable. It decouples them,  
while still not making it impossible to assimilate them to each other  
when necessary.

David

> **···**
>
> On Mon, 11 Jul 2005, Daniel Brockman wrote:
> 
> --  
> David A. Black  
> dblack@wobblini.net

[Previous page](https://rubytalk.org/t/strings-vs-arrays/19440.md?page=1)
