Wednesday, February 20, 2013

If you want to be a data scientist, you should know these 67 questions.

Recently, I read a blog by Vincent Granville on Data Science Central. If you would like to apply for a data scientist job or even prepare for this kind of job, you may try to answer these 66 open-ended questions.

66 job interview questions for data scientists

Here are some questions from it:

What is the curse of big data? (answer)

Examples where mapreduce does not work? Examples where it works very well? What are the security issues involved with the cloud? What do you think of EMC's solution offering an hybrid approach - both internal and external cloud - to mitigate the risks and offer other advantages (which ones)?  (answer)

I would like to add a question (not Why not 67 questions?):

Why do we need more data scientists? (answer)

Tuesday, February 12, 2013

Big Data in Financial Service Industry

In order to get senior management's buy-in on Big Data, you will have to show them some use cases.

Let's start from the financial service industry including the banks and others.

From Oracle:

This Oracle White Paper briefly talks about Oracle Big Data technology and several use cases in the financial services industry.

Financial Services Data Management:Big Data Technology in Financial Services

From IBM:

IBM solutions for big data provides banks with an integrated and scalable set of cost-effective, high-performance tools that support the rapid ingestion of important customer data from a variety of sources and the fast analysis of large volumes of data at transactional, product or enterprise levels.

See the link from IBM website: Deriving Business Insight from Big Data in Banking

And White Paper: IBM Information Agenda for Banking - Financial Crisis and Integrated Risk Management for Financial Institutions

From IDC:

The document is not free. You will have to  pay US$1,000 to get it.

Big Data - Use Cases in Financial ServicesPrice: US $1,000

Author: Michael Versace

Insights Presentation
July, 2012  -  Doc # FIN236035
Number of Pages: 18
Abstract
Data is the currency of competition in financial service. The effective use of data and information is the foundation upon which firms compete. Services are wrapped around data to differentiate products and services. For example, knowing which customers represent the best credit revenue and profitability opportunity to a bank is a question that only data and analysis can answer.
As an extension, IDC Financial Insights believes that Big Data and business analytics can quickly deliver competitive advantage for those firms that effectively harness and leverage the trend.. In this IDC Financial Insights presentation, we describe some of the drivers behind big data with examples for how big data technologies are being applied against some demanding business imperatives in the financial markets today. The presentation concludes with Essential Questions and Guidance to practitioners.


Sunday, February 10, 2013

The History of Big Data - 2

In my blog "The History of Big Data", it says the name "big data" originated as a tag for a class of technology with roots in high-performance computing.

After reading NYTimes article "The Origins of ‘Big Data’: An Etymological Detective Story" ,  I found out the origins of Big Data might not be different. 

In the article, it mentioned Francis X. Diebold, an economist at the University of Pennsylvania and his most recent paper. In the paper it concludes: “The term Big Data, which spans computer science and statistics/econometrics, probably originated in the lunch-table conversations at Silicon Graphics in the mid-1990s, in which John Mashey figured prominently.”

Wednesday, February 6, 2013

Big Data skills

To get into the field of Big Data, lots of people especially IT professionals are wondering what kinds of skills are required.

Here are some skills you should have or plan to have:



It will take time to learn and explore. But all the above skills will help you build your Big Data career path such as Data Scientist.

Here are some articles for your reference:

"Big data analytics is sometimes sold as a boon for IT workers, with analyst house Gartner predicting that within three years there will be 4.4 million staff working on big data projects. "

"The U.S. faces a substantial shortage of workers with data science skills, according to a much-talked about report published last year by consulting firm McKinsey and Company. The report predicted that by 2018 the country will lack 1.5 million analysts who can make strategic decisions using big data and between 140,000 to 190,000 workers with the proper data-processing technology skills."

"Regardless if they are called Data Scientists or Data Analysts, Data geeks need to be more in control of their destiny. "

Saturday, February 2, 2013

Big Data University

You may be wondering where you should start your Big Data learning journey. After a bit research, I found Big Data University is a good place to try.

Big Data University is an online educational site run by new and experienced Hadoop, Big Data and DB2 users who want to learn, contribute with course materials, or look for job opportunities. And it is hosted on the Cloud and using Moodle 2 course management system enabled to run on DB2. It is in the Beta stage.

The site includes free and fee-based courses delivered by experienced professionals and teachers.
When I saw DB2 but not other databases (including open source ones), I guess this site is either sponsored by IBM or run by IBM product lovers. Anyway, it is no harmful for you to learn Big Data.


In order to study in this "university", you should register by either using your Google, Facebook, Yahoo or ChannelDB2 account or creating your Big Data University account. Most of IT or data professionals should already have at least an account from Google, Facebook or Yahoo. If you don't use DB2, you might not even know ChannelDB2.


According to the site statistics, there are 63339 registered students (as of today - Feb.2, 2013 - not sure if it publishes the latest number). If you put this number under the perspective of real universities,  it is about 3 times size of Harvard (about 20,000 students)  or Stanford (about 18,000).


So, you want to join?

Sunday, January 27, 2013

Big Data Use Case #2 - Netflix

Just last week, Netflix stock soared after it fourth-quarter results top forecasts. On Jan.24, shares of Netflix rose $43.60 to $146.86 on Nasdaq, their highest level since September 2011.

Also in its earnings report, the company predicted it will add as many as 2.1 million U.S. streaming members in the first quarter, more than it gained during the first three months of last year.

How will Netflix to attract new subscribers? Although the company didn't tell, analyzing the Big Data should be one of the techniques. The people who has been following Netflix should know the Netflix Prize contest. The following provides a bit detail:

Netflix is all about connecting people to the movies they love. To help customers find those movies, we’ve developed our world-class movie recommendation system: CinematchSM. Its job is to predict whether someone will enjoy a movie based on how much they liked or disliked other movies. We use those predictions to make personal movie recommendations based on each customer’s unique tastes. And while Cinematch is doing pretty well, it can always be made better.

Now there are a lot of interesting alternative approaches to how Cinematch works that we haven’t tried. Some are described in the literature, some aren’t. We’re curious whether any of these can beat Cinematch by making better predictions. Because, frankly, if there is a much better approach it could make a big difference to our customers and our business.

So, we thought we’d make a contest out of finding the answer. It’s “easy” really. We provide you with a lot of anonymous rating data, and a prediction accuracy bar that is 10% better than what Cinematch can do on the same training data set. (Accuracy is a measurement of how closely predicted ratings of movies match subsequent actual ratings.) If you develop a system that we judge most beats that bar on the qualifying test set we provide, you get serious money and the bragging rights. But (and you knew there would be a catch, right?) only if you share your method with us and describe to the world how you did it and why it works.

Serious money demands a serious bar. We suspect the 10% improvement is pretty tough, but we also think there is a good chance it can be achieved. It may take months; it might take years. So to keep things interesting, in addition to the Grand Prize, we’re also offering a $50,000 Progress Prize each year the contest runs. It goes to the team whose system we judge shows the most improvement over the previous year’s best accuracy bar on the same qualifying test set. No improvement, no prize. And like the Grand Prize, to win you’ll need to share your method with us and describe it for the world.

According to the company blog,  Netflix announced the $1M Grand Prize winner of the Netflix Prize contest as team BellKor’s Pragmatic Chaos for their verified submission on July 26, 2009 at 18:18:28 UTC, achieving the winning RMSE of 0.8567 on the test subset.  This represents a 10.06% improvement over Cinematch’s score on the test subset at the start of the contest.

To know how much Cinematch has contributed to Netflix's financial result, it will need another project to make the calculation. One thing for sure, the company should collect more data from its subscribers not only from its business but also from other social source. The more the data they get, the better the recommendation they should provide, the larger the revenue they should make.

Friday, January 25, 2013

Big Data Use Case #1 - NBA

Have you ever heard of a company named Ayasdi? I didn't know this name until I recently Sarah Reedy's blog Ideas Watch: Ayasdi Gives Big-Data a Name.

In her blog, she talked about Ayasdi just got $10.25 million in Series A funding. For what? Ayasdi's cloud-based Insight Discovery Platform uses distributed computing, machine learning, and user-experience technologies to take all the guess work out of massive data sets. In company's own website, it says "Solving Today’s Biggest Problems Requires an Entirely New Approach to Data" and "A New Way to Discover Insights Leading to Breakthrough Outcomes".

I was amazed by the following picture named "Big-Data Basketball". If I didn't read the note under the picture, I thought it was about the new discovered galaxies by NASA or some new genetic maps found by scientists. It is actually a topological similarity network of 452 NBA players during the 2010-2011 season. Ayasi used its software to discover patterns from those NBA players' data and broke down the player into 13 classifications beyond the 5 normal positions on the court ( point guard, shooting guard, small forward, power forward and center).
 

Then what? The result from the analysis could change how coaches and general managers think about the roles their players fill and help team win more games. Also, the analysis could help team find good players and potential good players. In other words, the software makes the Big Data create value (money).  You can get more detail from the WIRED magazine article "Analytics Reveal 13 New Basketball Positions".

This use case also tells that this valuable analysis of big data was not done by those large companies like IBM, Oracle and Microsoft, but a startup.

Big Data provides huge opportunities to the startup companies.