Showing posts with label hadoop. Show all posts
Showing posts with label hadoop. Show all posts

Tuesday, July 29, 2014

Hadoop at Pinterest



Big data plays a big role at Pinterest. With more than 30 billion Pins in the system, we’re building the most comprehensive collection of interests online. One of the challenges associated with building a personalized discovery engine is scaling our data infrastructure to traverse the interest graph to extract context and intent for each Pin.
We currently log 20 terabytes of new data each day, and have around 10 petabytes of data in S3. We use Hadoop to process this data, which enables us to put the most relevant and recent content in front of Pinners through features such as Related Pins, Guided Search, and image processing. It also powers thousands of daily metrics and allows us to put every user-facing change through rigorous experimentation and analysis.


Building a self-serve platform for Hadoop

Though Hadoop is a powerful processing and storage system, it’s not a plug and play technology. Because it doesn’t have cloud or elastic computing, or non-technical users in mind, its original design falls short as a self-serve platform. Fortunately there are many Hadoop libraries/applications and service providers that offer solutions to these limitations. Before choosing from these solutions, we mapped out our Hadoop setup requirements.
1. Isolated multitenancy
2. Elasticity
3. Multi-cluster support
4. Support for ephemeral clusters
5. Easy software package deployment
6. Shared data store
7. Access control layer 
Read all 7 items more on engineering.pinterest.com

Friday, August 10, 2012

Bu hafta ilgimi ceken 3 link

  1. What can “Hadoop” do for genomics?
  2. Scoop.it - Epeydir gozatmayi erteledigim bir icerik kesif araci (benim icin sadece kesfetme araci)
  3. MongoDb 2.2 RC0 - Gecen hafta beta'dan cikti ama buyuk veriyle bu hafta denedim. Ozellikle TTL collections mevzusu artik yerine oturuyor. Ayrica  hakkinda bir kac notum var surada.




Tuesday, February 14, 2012

David Zuelke : Large-Scale Data Processing With Hadoop and PHP

Kabaca map reduce nedir sorusunu sorduysaniz once bu linkBu sunum da David Zuelke'nin son 6 aydir her konferansta sundugu ve benim de cok begendigim bir cirpida Hadoop-PHP sunumu : Large-Scale Data Processing With Hadoop and PHP (PHPBNL2012 2012-01-27)
View more presentations from David Zuelke