book

Q: What is big data?

  • Find FOUR really big datasets. They should be sufficiently different. Cite your sources.
  • Discuss and determine a ranking in terms of "bigness".
  • Fill the template below. Replace all (( )) with your answers.

Rank 1: amazon.com

The entirety of data that is collected/used/maintained by Amazon.com. This includes product inventory, user data, puchase history, Amazon Web Services, etc.

Amazon's data bigger than twitter's data because Amazon has a much larger variety of data, and in some instances, more complex data than twitter.

Rank 2: twitter.com

All tweet data contained by twitter.

Twitter's data is begger than the data of GoogleMaps, because Twitter's data is constantly growing, while the data contained in GoogleMaps is relatively unchanging and does not grow larger over time.

Rank 3: maps.google.com

The map data owned by Google; this includes both maps and images found in GoogleEarth.

Google Map data is larger than a human genome because the data in GoogleMaps is primarily image data, which can be difficult to work with, and a human genome (i.e. nucleotide sequences) is relativley simple to store and manage with computers.

Rank 4: the human genome (example: 1000genomes.org/data)

The data contained in a single human genome.