# FasterCSV - Merge CSV

**URL:** https://rubytalk.org/t/fastercsv-merge-csv/59058
**Category:** ruby-talk
**Created:** [2 July 2010 07:35 UTC](https://rubytalk.org/t/fastercsv-merge-csv/59058 "2010-07-02T07:35:39Z")
**Posts on this page:** 7
**Page:** 1

<div class="post-metadata">

### Author: ![Christian\_Smith](https://avatars.discourse-cdn.com/v4/letter/c/22d042/32.png) [@Christian\_Smith](https://rubytalk.org/u/Christian_Smith)
#### Post date: [2 July 2010 07:35 UTC](https://rubytalk.org/t/fastercsv-merge-csv/59058/1 "2010-07-02T07:35:39Z")

</div>

I have 3 CSVs with the same content with say 10 rows. There is a slight  
variation to 1 column in the data it contains.

csv1 - 20 lines 10 cols  
csv2 - 52 lines 10 cols  
csv3 - 24 lines 10 cols

How can I merge all 3 csvs into 1 csv using fastercsv so I have

csv4 96 lines 10 cols

Thanks!

Seed

> **···**
>
> --  
> Posted via [http://www.ruby-forum.com/](http://www.ruby-forum.com/).

---

<div class="post-metadata">

### Author: ![Brian\_Candler](https://avatars.discourse-cdn.com/v4/letter/b/5f9b8f/32.png) [@Brian\_Candler](https://rubytalk.org/u/Brian_Candler)
#### Post date: [2 July 2010 11:11 UTC](https://rubytalk.org/t/fastercsv-merge-csv/59058/2 "2010-07-02T11:11:24Z")

</div>

Christian Smith wrote:

> I have 3 CSVs with the same content with say 10 rows. There is a slight  
> variation to 1 column in the data it contains.
> 
> csv1 - 20 lines 10 cols  
> csv2 - 52 lines 10 cols  
> csv3 - 24 lines 10 cols
> 
> How can I merge all 3 csvs into 1 csv using fastercsv so I have
> 
> csv4 96 lines 10 cols
> 
> Thanks!
> 
> Seed

Why use fastercsv?  
cat csv1 csv2 csv3 \>csv4  
would meet your requirement.

But if you want to use fastercsv, then open each file in turn, read it  
line at a time, and output the line you just read.

> **···**
>
> --  
> Posted via [http://www.ruby-forum.com/\](http://www.ruby-forum.com/%5C).

---

<div class="post-metadata">

### Author: ![Rob\_Biedenharn1](https://avatars.discourse-cdn.com/v4/letter/r/57b2e6/32.png) [@Rob\_Biedenharn1](https://rubytalk.org/u/Rob_Biedenharn1)
#### Post date: [2 July 2010 13:54 UTC](https://rubytalk.org/t/fastercsv-merge-csv/59058/3 "2010-07-02T13:54:47Z")

</div>

> Christian Smith wrote:
> 
> > I have 3 CSVs with the same content with say 10 rows. There is a slight  
> > variation to 1 column in the data it contains.
> > 
> > csv1 - 20 lines 10 cols  
> > csv2 - 52 lines 10 cols  
> > csv3 - 24 lines 10 cols
> > 
> > How can I merge all 3 csvs into 1 csv using fastercsv so I have
> > 
> > csv4 96 lines 10 cols
> > 
> > Thanks!
> > 
> > Seed
> 
> Why use fastercsv?  
> cat csv1 csv2 csv3 \>csv4  
> would meet your requirement.

except that you'd have headers from csv2 and csv3 (but perhaps your line counts imply no headers?)

> But if you want to use fastercsv, then open each file in turn, read it  
> line at a time, and output the line you just read.  
> --

If the files are small-ish, you can avoid a chicken-and-egg problem of the headers by reading all the input files (saving the headers from the first), then writing it all out from memory.

-Rob

Rob Biedenharn   
Rob@AgileConsultingLLC.com [http://AgileConsultingLLC.com/](http://AgileConsultingLLC.com/)  
rab@GaslightSoftware.com [http://GaslightSoftware.com/](http://GaslightSoftware.com/)

> **···**
>
> On Jul 2, 2010, at 7:11 AM, Brian Candler wrote:

---

<div class="post-metadata">

### Author: ![Christian\_Smith](https://avatars.discourse-cdn.com/v4/letter/c/22d042/32.png) [@Christian\_Smith](https://rubytalk.org/u/Christian_Smith)
#### Post date: [2 July 2010 17:15 UTC](https://rubytalk.org/t/fastercsv-merge-csv/59058/4 "2010-07-02T17:15:06Z")

</div>

Rob Biedenharn wrote:

> > > csv4 96 lines 10 cols
> > > 
> > > Thanks!
> > > 
> > > Seed
> > 
> > Why use fastercsv?  
> > cat csv1 csv2 csv3 \>csv4  
> > would meet your requirement.
> 
> except that you'd have headers from csv2 and csv3 (but perhaps your  
> line counts imply no headers?)
> 
> > But if you want to use fastercsv, then open each file in turn, read it  
> > line at a time, and output the line you just read.  
> > --
> 
> If the files are small-ish, you can avoid a chicken-and-egg problem of  
> the headers by reading all the input files (saving the headers from  
> the first), then writing it all out from memory.
> 
> -Rob
> 
> Rob Biedenharn  
> Rob@AgileConsultingLLC.com [http://AgileConsultingLLC.com/](http://AgileConsultingLLC.com/)  
> rab@GaslightSoftware.com [http://GaslightSoftware.com/](http://GaslightSoftware.com/)

If the files are small-ish, you can avoid a chicken-and-egg problem of  
the headers by reading all the input files (saving the headers from  
the first), then writing it all out from memory.

The files aren't smallish but memory isn't an issue. I would love to be  
able to do this. I am able to read the 3 files into an array but it's  
parsing them back into 1 csv I am having trouble with. I would assume  
this would be a lot faster than a line read\>write approach.

> **···**
>
> > On Jul 2, 2010, at 7:11 AM, Brian Candler wrote:
> 
> --  
> Posted via [http://www.ruby-forum.com/\](http://www.ruby-forum.com/%5C).

---

<div class="post-metadata">

### Author: ![Reid\_Thompson1](https://avatars.discourse-cdn.com/v4/letter/r/6de8d8/32.png) [@Reid\_Thompson1](https://rubytalk.org/u/Reid_Thompson1)
#### Post date: [2 July 2010 17:31 UTC](https://rubytalk.org/t/fastercsv-merge-csv/59058/5 "2010-07-02T17:31:18Z")

</div>

just cat and grep out the header lines

cat csv\* |grep -v string-portion-unique-to-headers \> full.csv

If you want to head a header row, then

cat csv\* | grep string-portion-unique-to-headers |sort | uniq \> full.csv  
cat csv\* | grep -v string-portion-unique-to-headers \>\> full.csv

> **···**
>
> On Sat, Jul 03, 2010 at 02:15:06AM +0900, Christian Smith wrote:
> 
> > Rob Biedenharn wrote:  
> > \> On Jul 2, 2010, at 7:11 AM, Brian Candler wrote:  
> > \>\>\>  
> > \>\>\> csv4 96 lines 10 cols  
> > \>\>\>  
> > \>\>\> Thanks!  
> > \>\>\>  
> > \>\>\> Seed  
> > \>\>  
> > \>\> Why use fastercsv?  
> > \>\> cat csv1 csv2 csv3 \>csv4  
> > \>\> would meet your requirement.  
> > \>  
> > \> except that you'd have headers from csv2 and csv3 (but perhaps your  
> > \> line counts imply no headers?)  
> > \>  
> > \>\>  
> > \>\> But if you want to use fastercsv, then open each file in turn, read it  
> > \>\> line at a time, and output the line you just read.  
> > \>\> --  
> > \>  
> > \> If the files are small-ish, you can avoid a chicken-and-egg problem of  
> > \> the headers by reading all the input files (saving the headers from  
> > \> the first), then writing it all out from memory.  
> > \>  
> > \> -Rob  
> > \>  
> > \> Rob Biedenharn  
> > \> Rob@AgileConsultingLLC.com [http://AgileConsultingLLC.com/](http://AgileConsultingLLC.com/)  
> > \> rab@GaslightSoftware.com [http://GaslightSoftware.com/](http://GaslightSoftware.com/)
> > 
> > If the files are small-ish, you can avoid a chicken-and-egg problem of  
> > the headers by reading all the input files (saving the headers from  
> > the first), then writing it all out from memory.
> > 
> > The files aren't smallish but memory isn't an issue. I would love to be  
> > able to do this. I am able to read the 3 files into an array but it's  
> > parsing them back into 1 csv I am having trouble with. I would assume  
> > this would be a lot faster than a line read\>write approach.
> > 
> > --  
> > Posted via [http://www.ruby-forum.com/\](http://www.ruby-forum.com/%5C).

---

<div class="post-metadata">

### Author: ![Rob\_Biedenharn1](https://avatars.discourse-cdn.com/v4/letter/r/57b2e6/32.png) [@Rob\_Biedenharn1](https://rubytalk.org/u/Rob_Biedenharn1)
#### Post date: [2 July 2010 19:27 UTC](https://rubytalk.org/t/fastercsv-merge-csv/59058/6 "2010-07-02T19:27:40Z")

</div>

OK, let's read them all in and then write out one file...

headers = nil  
all\_rows =   
input\_files.each do |input\_file|  
&nbsp;&nbsp;csv = FasterCSV.table(input\_file, :headers =\> true)  
&nbsp;&nbsp;in\_headers, \*in\_rows = csv.to\_a  
&nbsp;&nbsp;headers ||= in\_headers  
&nbsp;&nbsp;all\_rows.concat(in\_rows)  
end  
FasterCSV.open(output\_file, 'w') do |csv|  
&nbsp;&nbsp;&nbsp;csv \<\< headers  
&nbsp;&nbsp;&nbsp;all\_rows.each {|row| csv \<\< row }  
end

The full example is at:  
&nbsp;&nbsp;&nbsp;[combined.csv · GitHub](http://gist.github.com/461784)

The details may have to change a bit depending on your circumstances, but the general idea is sound.

-Rob

Rob Biedenharn   
Rob@AgileConsultingLLC.com [http://AgileConsultingLLC.com/](http://AgileConsultingLLC.com/)  
rab@GaslightSoftware.com [http://GaslightSoftware.com/](http://GaslightSoftware.com/)

> **···**
>
> On Jul 2, 2010, at 1:15 PM, Christian Smith wrote:
> 
> > Rob Biedenharn wrote:
> > 
> > > On Jul 2, 2010, at 7:11 AM, Brian Candler wrote:
> > > 
> > > > > csv4 96 lines 10 cols
> > > > > 
> > > > > Thanks!
> > > > > 
> > > > > Seed
> > > > 
> > > > Why use fastercsv?  
> > > > cat csv1 csv2 csv3 \>csv4  
> > > > would meet your requirement.
> > > 
> > > except that you'd have headers from csv2 and csv3 (but perhaps your  
> > > line counts imply no headers?)
> > > 
> > > > But if you want to use fastercsv, then open each file in turn, read it  
> > > > line at a time, and output the line you just read.  
> > > > --
> > > 
> > > If the files are small-ish, you can avoid a chicken-and-egg problem of  
> > > the headers by reading all the input files (saving the headers from  
> > > the first), then writing it all out from memory.
> > > 
> > > -Rob
> > > 
> > > Rob Biedenharn  
> > > Rob@AgileConsultingLLC.com [http://AgileConsultingLLC.com/](http://AgileConsultingLLC.com/)  
> > > rab@GaslightSoftware.com [http://GaslightSoftware.com/](http://GaslightSoftware.com/)
> > 
> > If the files are small-ish, you can avoid a chicken-and-egg problem of  
> > the headers by reading all the input files (saving the headers from  
> > the first), then writing it all out from memory.
> > 
> > The files aren't smallish but memory isn't an issue. I would love to be  
> > able to do this. I am able to read the 3 files into an array but it's  
> > parsing them back into 1 csv I am having trouble with. I would assume  
> > this would be a lot faster than a line read\>write approach.

---

<div class="post-metadata">

### Author: ![Brabuhr](https://avatars.discourse-cdn.com/v4/letter/b/919ad9/32.png) [@Brabuhr](https://rubytalk.org/u/Brabuhr)
#### Post date: [2 July 2010 20:19 UTC](https://rubytalk.org/t/fastercsv-merge-csv/59058/7 "2010-07-02T20:19:00Z")

</div>

FasterCSV.open(output\_file, 'w') do |ocsv|  
&nbsp;&nbsp;input\_files.each\_with\_index do |input\_file, i|  
&nbsp;&nbsp;&nbsp;&nbsp;FasterCSV.foreach(input\_file, :headers =\> true, :return\_headers =\>  
true) do |row|  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;next if i \> 0 and row.header\_row?  
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ocsv \<\< row  
&nbsp;&nbsp;&nbsp;&nbsp;end  
&nbsp;&nbsp;end  
end

> **···**
>
> On Fri, Jul 2, 2010 at 3:27 PM, Rob Biedenharn \<Rob@agileconsultingllc.com\> wrote:
> 
> > OK, let's read them all in and then write out one file...
> > 
> > headers = nil  
> > all\_rows =   
> > input\_files.each do |input\_file|  
> > csv = FasterCSV.table(input\_file, :headers =\> true)  
> > in\_headers, \*in\_rows = csv.to\_a  
> > headers ||= in\_headers  
> > all\_rows.concat(in\_rows)  
> > end  
> > FasterCSV.open(output\_file, 'w') do |csv|  
> > csv \<\< headers  
> > all\_rows.each {|row| csv \<\< row }  
> > end
