Thursday, July 9, 2015

Selling analytics (HR Analytics

Predictive Analytic in HR 

Business Objective: To clearly demonstrate the interaction of business objective and work force strategies to determine a full picture of likely outcomes

Based on research, Global organization with workforce analytics and workforce planning outperform all other organisation by 30%

Outlining the tree step Process for Predictive Analytic in HR Function
1. Hindsight: Gather data by Reporting
2. Insights: Make sense of data by Analysis and monitoring
3. Foresight: Develop predictive Models

KPI/Matrices
What is generally measured:
- Employee Engagement
- Performance Ratings
- Retention/Turnover
- % of employee with dev plan
- Reediness for jobs
- Internal hire %age
- Diversity of workforce
- Level of expertise/competence

HR Analytics:
What could be measured:
- Recruitment
- Retention
- Performance and Career Management
- Training
- Compensation and Benefits
- Workfore
- Organization and effectiveness

Benefits
Turnover modeling: Predicting future turnover in business units in specific functions, geographies by looking at factor such as commute time, time since last role change and performance over time
Targeted retention: Find out high risk of churn in future and focus on retention of few critical people
3. Risk Management: Profiling of candidate with high risk of leaving prematurely or those performing below standard
4. Talent Forecasting: To predict which new hire, based on their profile, are likely to be high flier and then moving to them on fast track programs.

Overall Return on the Investment
- Retention of key performance and their associated customer and revenue
- Reduced compensation over payment
- Improved HR staff  productivity  resulting in the need of fewer HR Staf
- Reduced risk of litigation due to non-compliance




Thursday, July 2, 2015

Tree map and Geospatial map

The term mashup started in the music world but  adopted rapidly in the world of web that means an application that combines data from different sources into a whole new application

Geographic data is one of the most common types of data available. Dealing with geographic data without a map like going into the mountain without, well, a map
Geospatial analysis help in identifying "hot spots" where disease or problems are occurring.
Please refer the below article how dominos is leveraging geospatial analytics to know their customer and stores performance.
http://www.zdnet.com/article/dominos-pizza-gets-customer-specific-using-geospatial-analytics/

Treemapping is a method for displaying hierarchical data by using nested rectangle
Treemap display hierarchical(tree-structured) data as a set of nested rectangles. Each branch of the tree is given a rectangle, which is then tiled with smaller rectangle representing sub branches.
Common use case we are seeing these days in analyzing disk space. There is utility called treesize being used in monitoring disk space on servers

Here is very interesting software developed by MIT student which is based on the treemap.

http://pantheon.media.mit.edu/treemap/country_exports/IN/all/-4000/2010/H15/pantheon

For further indepth reading on treemap from NorthWestern University
http://www.cs.uic.edu/~wilkinson/Publications/c&rtrees.pdf


Reference:
https://en.wikipedia.org/wiki/Treemapping
http://searchbusinessanalytics.techtarget.com/news/1507131/Data-mashups-meet-business-intelligence-Bashups-explained


Predicting Wine Quality Analytically


Abstract: Wine industry shows growth in overall consumption of wine. Price of wine depend on two critical factor
- Wine appreciation by wine tester
- Certification and quality assessment in the physicochemical test
We have two dataset red wine and white wine.
We have done exploratory data analysis using standard function in R. During the analysis, we have identified the outliers in different variables using box plot. Also using cor() R function tried to understand the correlation between Quality and rest of variables.
Finally devised model using Liner regression technique to predict the quality of wine
Red wine data used to depict the picture. Same steps can be applied on white wine data

Project Goal:
- Explore the data in dataset and be able to list all the standard summary statistics
- investigate distribution of the variables graphically to determine the outliers
- devise method to handle outlier
- investigate correlation between quality and remaining properties

- suggest methods for the final “Quality” determination



Looking into dataset(Red Wine)
> red.wine.data <-read.delim(file.choose(), header=T)
> dim(red.wine.data)
[1] 1599   12
> names(red.wine.data)
 [1] "fixed.acidity"        "volatile.acidity"     "citric.acid"          "residual.sugar"     
 [5] "chlorides"            "free.sulfur.dioxide"  "total.sulfur.dioxide" "density"            
 [9] "pH"                   "sulphates"            "alcohol"              "quality"
> str(red.wine.data)
'data.frame':       1599 obs. of  12 variables:
 $ fixed.acidity       : num  7.4 7.8 7.8 11.2 7.4 7.4 7.9 7.3 7.8 7.5 ...
 $ volatile.acidity    : num  0.7 0.88 0.76 0.28 0.7 0.66 0.6 0.65 0.58 0.5 ...
 $ citric.acid         : num  0 0 0.04 0.56 0 0 0.06 0 0.02 0.36 ...
 $ residual.sugar      : num  1.9 2.6 2.3 1.9 1.9 1.8 1.6 1.2 2 6.1 ...
 $ chlorides           : num  0.076 0.098 0.092 0.075 0.076 0.075 0.069 0.065 0.073 0.071 ...
 $ free.sulfur.dioxide : num  11 25 15 17 11 13 15 15 9 17 ...
 $ total.sulfur.dioxide: num  34 67 54 60 34 40 59 21 18 102 ...
 $ density             : num  0.998 0.997 0.997 0.998 0.998 ...
 $ pH                  : num  3.51 3.2 3.26 3.16 3.51 3.51 3.3 3.39 3.36 3.35 ...
 $ sulphates           : num  0.56 0.68 0.65 0.58 0.56 0.56 0.46 0.47 0.57 0.8 ...
 $ alcohol             : num  9.4 9.8 9.8 9.8 9.4 9.4 9.4 10 9.5 10.5 ...
 $ quality             : int  5 5 5 6 5 5 5 7 7 5 ...
> attributes(red.wine.data)
$names
 [1] "fixed.acidity"        "volatile.acidity"     "citric.acid"          "residual.sugar"     
 [5] "chlorides"            "free.sulfur.dioxide"  "total.sulfur.dioxide" "density"            
 [9] "pH"                   "sulphates"            "alcohol"              "quality"            
$class
[1] "data.frame"
> summary(red.wine.data)
 fixed.acidity   volatile.acidity  citric.acid    residual.sugar     chlorides     
 Min.   : 4.60   Min.   :0.1200   Min.   :0.000   Min.   : 0.900   Min.   :0.01200 
 1st Qu.: 7.10   1st Qu.:0.3900   1st Qu.:0.090   1st Qu.: 1.900   1st Qu.:0.07000 
 Median : 7.90   Median :0.5200   Median :0.260   Median : 2.200   Median :0.07900 
 Mean   : 8.32   Mean   :0.5278   Mean   :0.271   Mean   : 2.539   Mean   :0.08747 
 3rd Qu.: 9.20   3rd Qu.:0.6400   3rd Qu.:0.420   3rd Qu.: 2.600   3rd Qu.:0.09000 
 Max.   :15.90   Max.   :1.5800   Max.   :1.000   Max.   :15.500   Max.   :0.61100 
 free.sulfur.dioxide total.sulfur.dioxide    density             pH          sulphates    
 Min.   : 1.00       Min.   :  6.00       Min.   :0.9901   Min.   :2.740   Min.   :0.3300 
 1st Qu.: 7.00       1st Qu.: 22.00       1st Qu.:0.9956   1st Qu.:3.210   1st Qu.:0.5500 
 Median :14.00       Median : 38.00       Median :0.9968   Median :3.310   Median :0.6200 
 Mean   :15.87       Mean   : 46.47       Mean   :0.9967   Mean   :3.311   Mean   :0.6581 
 3rd Qu.:21.00       3rd Qu.: 62.00       3rd Qu.:0.9978   3rd Qu.:3.400   3rd Qu.:0.7300 
 Max.   :72.00       Max.   :289.00       Max.   :1.0037   Max.   :4.010   Max.   :2.0000 
    alcohol         quality    
 Min.   : 8.40   Min.   :3.000 
 1st Qu.: 9.50   1st Qu.:5.000 
 Median :10.20   Median :6.000 
 Mean   :10.42   Mean   :5.636
 3rd Qu.:11.10   3rd Qu.:6.000 
 Max.   :14.90   Max.   :8.000

Identifying outliers

> fa<-red.wine.data$fixed.acidity
> boxplot(fa)





Handling outliers

Removed the outlier values and stored data in new data frame called red.wine

> red.wine<-subset(red.wine.data, fa< 12 & va<1 & ca<.8 & rs< 5 & cl<.2& fsd < 50 & tsd<150 & d<1 & ph<3.5 & sul<1.5 & al< 14 & ql<9)

> summary(red.wine)

fixed.acidity    volatile.acidity  citric.acid     residual.sugar    chlorides     
 Min.   : 5.000   Min.   :0.1200   Min.   :0.0000   Min.   :0.900   Min.   :0.01200 
 1st Qu.: 7.200   1st Qu.:0.3800   1st Qu.:0.1100   1st Qu.:1.900   1st Qu.:0.07000 
 Median : 8.000   Median :0.5100   Median :0.2600   Median :2.100   Median :0.07900 
 Mean   : 8.276   Mean   :0.5093   Mean   :0.2657   Mean   :2.243   Mean   :0.08112 
 3rd Qu.: 9.100   3rd Qu.:0.6200   3rd Qu.:0.4000   3rd Qu.:2.500   3rd Qu.:0.08900 
 Max.   :11.900   Max.   :0.9800   Max.   :0.7300   Max.   :4.800   Max.   :0.19400 
 free.sulfur.dioxide total.sulfur.dioxide    density             pH          sulphates    
 Min.   : 1.00       Min.   :  6.00       Min.   :0.9901   Min.   :2.870   Min.   :0.3300 
 1st Qu.: 7.00       1st Qu.: 21.50       1st Qu.:0.9956   1st Qu.:3.220   1st Qu.:0.5400 
 Median :13.00       Median : 37.00       Median :0.9966   Median :3.300   Median :0.6100 
 Mean   :15.39       Mean   : 44.58       Mean   :0.9965   Mean   :3.293   Mean   :0.6407 
 3rd Qu.:21.00       3rd Qu.: 59.00       3rd Qu.:0.9975   3rd Qu.:3.380   3rd Qu.:0.7100 
 Max.   :48.00       Max.   :149.00       Max.   :0.9998   Max.   :3.490   Max.   :1.3600 
    alcohol         quality    
 Min.   : 8.50   Min.   :3.000 
 1st Qu.: 9.50   1st Qu.:5.000 
 Median :10.10   Median :6.000 
 Mean   :10.37   Mean   :5.655 
 3rd Qu.:11.00   3rd Qu.:6.000 
 Max.   :13.60   Max.   :8.000  





Correlation: Principal component analysis

By doing principal component analysis and plotting, we can easily identify the principal components and their correlation.
#number of element
> temp_red.wine<-length(as.matrix(red.wine))/length(red.wine)
#PCA analysis
> pcx<-prcomp(red.wine, scale=TRUE)
#plotting using biplot

> biplot(pcx, xlab=rep('.', temp_red.wine))

Interesting about the plot is that judging by the first two principal components, a quality is very much correlated with alcohol content and sulphate


Predicting wine quality using liner regression line

> plot(ql~al, data=red.wine)
> mean(ql)
[1] 5.636023
> abline(h=mean(ql))
> model1=lm(ql~al, data=red.wine)
> model1

Call:
lm(formula = ql ~ al, data = red.wine)

Coefficients:
(Intercept)           al 
     1.8750       0.3608 

> abline(model1,col="red")
> plot(model1)




Wednesday, July 1, 2015

Difference between charts


Choosing right char for your data

There are found different types of categories for your data

Distribution: shows a collection of related/unrelated information to see how it correlates, if at all

Composition: Collecting different types of data that make up a whole and displaying them together

Comparison: Sets variable apart from the each other and shows how they interact

Relationship: This type of data tries to show a correlation between two or more variables


Histograms are used to show distribution of variables while bar chats are used to compare variables

Histograms plot quantitative data with ranges of the data grouped into bins or intervals while bar charts plot categorical data

There is no space between bars in histograms where is there is space between bars in bar chart


In bar graphs are usually used to display "categorical data", that is data that fits into categories.
Histograms on the other hand are usually used to present "continuous data"

Example:
Bar Chart: Average Per Capita income of metro cities Delhi, Mumbai, Kolkata, Chennai
Histogram: Average Per Capita income of five age groups(25-35, 35-45, 45,55, 55-65)



Histogram is subclass of bar graph
https://www.youtube.com/watch?v=F0zxwK7-OeY

A Histogram is NOT a Bar Chart
http://www.forbes.com/sites/naomirobbins/2012/01/04/a-histogram-is-not-a-bar-chart/



Reference:
Tree Structured Data Analysis: AID, CHAID and CART
http://www.cs.uic.edu/~wilkinson/Publications/c&rtrees.pdf

Sunday, June 21, 2015

Descriptive Vs. Predictive Vs. Prescriptive

Descriptive Analytics (EDA)
The purpose of descriptive analytics is simply to summarize and tell you what happened. For example, number of post, mentions, fans, followers, page views, kudos, +1s, check-ins, pins, etc. There are literally thousands of these metrics – it’s pointless to list them – but they are all just simple event counters. Other descriptive analytics may be results of simple arithmetic operations, such as share of voice, average response time, % index, average number of replies per post, etc.

Predictive Analytics

The purpose of predictive analytics is NOT to tell you what will happen in the future. It cannot do that. In fact, no analytics can do that. Predictive analytics can only forecast what might happen in the future, because all predictive analytics are probabilistic in nature.

The essence of predictive analytics, in general, is that we use existing data to build a model. Then we use the model to predict data that doesn’t yet exist. So predictive analytics is all about using data you have to predict data that you don’t have.

Prescriptive Analytics
Presecriptinve analytics not only predicts a possible future , it predicts
multiple futures based on the decisions makes action. Therefore a prescriptive model is
by definition, also predictive. As such must be validated too.
A prescriptive model can be viewed as a combination of multiple predictive models
running in parallel one for each possible input action
The Goal of the most prescriptive analytics is to guide the decision maker so the decision he makes
will ultimately lead to target outcome

In prescriptive analytics, we also build a predictive model of data
The predictive model must have two more added component in order to be prescriptive
1. Actionable: The data consumers must be able to take action based on the predictive outcome of the model

2. Feedback System: The model must have feedback system that tracks the adjusted outomes
based on the action taken. This means the predictive model must be smart enough to learn the complex relationship between the users' action and the adjusted outcome through the feedback data


1. Descriptive Analytics: Compute descriptive statistics to summarize the data. The majority of social analytics fall in this category
2. Predictive Analytics: Build a statistical model that uses existing data to predict data that that we don’t have. Examples of predictive analytics include trend lines, influence scoring, sentiment analysis, etc.
3. Prescriptive Analytics: Build a prescriptive model that uses not only the existing data, but also the action and feedback data to guide the decision maker to a desired outcome. Because prescriptive models must be actionable and have a feedback data stream, social analytics are rarely prescriptive

Identifying outlier

Identifying outliers is called is separating signal from the noise
Business Analytics process goes with 3 steps:
- Framing
- Analysis
- Reporting
Each step pass through with multiple steps. And analysing has
- Data collection and understanding
- Data preparation
- Analysis and Model building
Under data preparation, we have filtering where we have to perform identifying outliers. Outliers are the data points which significantly differ from the sample data. Having such data points/outliers can result in abnormality in result
Hence finding out outliers is very important activity.

Let me introduce how to find the outlier..
Just take an example from the following sample data points(already organized in lowest to highest order)
10. 12, 13, 15, 16, 20, 35
Median: 7+1/2=4 and That is 15
Lower Quartile Q1: 12
Upper Quartile Q3: 20

Inter quartile Range: Q3  -  Q1  =  20 -12 = 8
Outlier (Lower Fence) < Lower Quartile(Q1) -  1.5(IQ) = 12 - 1.5(8) = 0
Outlier(Upper Fence ) > Upper Quartile(Q3) - 1.5(IQ) = 20 + 1,5(8) = 32
Hence, there is no value lower than 0 in above sample data but there is value greater than 32 which is 35. 
Accordingly in our sample data we have one data point that is 35 which can be considered as outlier.
Understand the importance of retaining outliers
While some outliers should be omitted from the data set because that can lead to bad result. However, we should be careful as that can be critical element in deriving the results.
Use a qualitative assessment in determining whether to "throw out" outliers. 
There are two graphical technique to identify the outliers that is scatter plot (Links to an external site.) and box plot (Links to an external site.) and analytical procedure Grubbs' Test  (Links to an external site.)for detecting outliers 

Filtering Outliers

Its General strategy to omit the outlier from the same. But, in every situation an outlier must not be candidate for omission.
As stated my my previous DQ 6.1 understanding the importance of retaining or throwing outlier is very important. 
We should use Qualitative Assessment  in determining whether to throw out the outlier.

Scientific experiment generally contain sensitive data and hence an outlier can give new trend/insight to experiment. 

Let me share the situation which can clarify that outlier is not always candidate for omission.

In a clinical trial conducted on an Hen in farm. Here is weekly(10 weeks) eggs produced by that Hen

10, 11, 14, 15, 12, 13, 30, 9, 16

In above data 30 eggs in the 7th week seem outlier. But we should not omit it because assuming its not due to error. This can result significant success in experiment. And the drug used in 7th week has given a better results. The experiment conducted on 7th week yielded successful results. 


Data Cleansing

Data cleaning also called data cleansing or scrubbing, deals with detecting and removing errors and inconsistencies from the data in order to improve the quality of data

Data quality discussion typically involve the following services
- Parsing
- Standardization
- Validation
- Verification
- Matching

There are tools in the market which worked on above process.  For example IBM Infoshpere Information Server for data Quality and Microsoft SQL Server 2016.
Microsoft 2016 includes Data Quality Services (DQS) which is a computer-assisted process that analyzes how data conforms to the knowledge in knowledge base DQS  categorized data data under following five tab
- Suggested
- New
- Invalid
- Corrected
- Correct

Consider the case of Bigdata where integrity aspects of data is very hot and now known as veracity of data.
Some case veracity of data can be maintained at the origin of data itself by enforcing field integrity.
For example:
Email field: email should be valid. It should contain validation on special character like @
Zip Code Field : Zip code field should have certain no. of digit. If  system is integrated with master data where each country code is store. We can enforce integrity of zip code by verifying zip code from master data
Mobile no Filed: Mobile no. can be validated by introducing OTP mechanism. This can insure the mobile no/contact information provided is correct

One of other approach in data cleansing may involve relationship integrity which is know as association rule. This concept was first introduced in market basket analysis

Above field integrity and relation ship integrity is mostly true in case of structured data collection
but what about in case of unstructured data?
Consider the case of social media based customer profiling. Big data center are coming with concept of Master Data Management (MDM). And linkage between Dataware house and MDM project is good starting point for providing cleaned data. For example: Consider the Master Data project centers around a person, then extracting life event over Twitter or Facebook, such as a change in relation ship status, birth announcement  and so on enriches that master information,










Wednesday, June 10, 2015

Structured - Unstructured data and DataWare House


DB4.1: Data types and large-scale analytics

3 unread replies.3 replies.



“Unstructured data refers to information that either does not have a pre-defined data model and/or is not organized in a predefined manner.”

In addition to social media there are many other common forms of unstructured data:
  • Word Doc’s, PDF’s and Other Text Files - Books, letters, other written documents, audio and video transcripts
  • Audio Files - Customer service recordings, voicemails, 911 phone calls
  • Presentations - PowerPoints, SlideShares
  • Videos - Police dash cam, personal video, YouTube uploads
  • Images - Pictures, illustrations, memes
  • Messaging - Instant messages, text messages
  • In all these instances, the data can provide compelling insights. Using the right tools, unstructured data can add a depth to data analysis that couldn’t be achieved otherwise.
Structured Data
Contrasting to unstructured data, structured data is data that can be easily organized. Regardless of its simplicity, most experts in today’s data industry estimate that structured data accounts for only 20% of the data available. It is clean, analytical and usually stored in databases.
- Sensory Data - GPS data, manufacturing sensors, medical devices
- Point-of-Sale Data - Credit card information, location of sale, product information
- Call Detail Records - Time of call, caller and recipient information
- Web Server Logs - Page requests, other server activity
- Input Data - Any data inputted into a computer: age, zip code, gender, etc.
The scale  and verity of data have permanently overwhelmed the ability to cost effectively  extract value using traditional platform
Choice of storage system and computing platform is depend on types of data and/or source of data and size of data. 
Today's' parallel computing platform is best suited platform to handler the speed and volume of data being produced.  There are 3 prominent computing platform options are available today.

- Cluster or Grids
- Massively Parallel processing(MPP)
- High performance computing (HPC)

If there is verity(structured and unstructured)  in data and coming for  different sources in big volume the Hadoop is best suited technology in these days.
Hadoop basically used as synonym for HDFS(Hadoop distributed File system) which comes with default parallel processing programming framework called "MapReduce"
MapReduce is fault tolerant parallel programming framework that was designed to harness distributing processing capabilities. MapReduce automatically divides the process workload into smaller workloads that are distributed.  
I would like to share few advantage which can help you to gauge the power of hadoop and help to decide the right storage platform.

Cost effective
- Its open source software which is govern by open source community under the APACHE License. It means no cost involved and we are open to manipulate and customized per our requirement.
- It work on commodity hardware hence its again cost effective 
Scalable
- Hadoop is a highly scalable storage platform, because it can store and distribute very large data sets across hundreds of inexpensive servers that operate in parallel
Flexible

Hadoop enables businesses to easily access new data sources and tap into different types of data (both structured and unstructured) to generate value from that data.
Fast

Hadoop’s unique storage method is based on a distributed file system that basically ‘maps’ data wherever it is located on a cluster. The tools for data processing are often on the same servers where the data is located, resulting in much faster data processing. If you’re dealing with large volumes of unstructured data, Hadoop is able to efficiently process terabytes of data in just minutes, and petabytes in hours.
Resilient to failure
A key advantage of using Hadoop is its fault tolerance. When data is sent to an individual node, that data is also replicated to other nodes in the cluster, which means that in the event of failure, there is another copy available for use.

HP Vertica:
The HP Vertica Analytics Platform was consciously designed with speed, scalability, simplicity, and openness at its core and architected to handle analytical workloads via a distributed compressed columnar architecture. HP Vertica provides blazingfast speed (queries run 50-1,000x faster), petabyte-scale (store 10-30x more data per server), and openness and simplicity (use
any BI/ETL tools, Hadoop, etc.) — all at 30% of the cost of traditional data warehouse solutions.
The Technology that Makes HP Vertica So Powerful
The HP Vertica Analytics Platform is a standards-based, relational database. It supports popular SQL, JDBC/ODBC. This allows users to preserve years of investment and training in these technologies because all popular SQL programming tools and languages work seamlessly. All popular BI and visualization tools are tightly integrated, such as Tableau, Microstrategy and others and so are all popular ETL tools like Informatica, Pentaho, and more. HP Vertica is optimized for large-scale analytics. It is uniquely designed using a memory-and-disk balanced distributed compressed columnar paradigm, which makes it exponentially faster than older techniques for modern data analytics workloads. Additionally, HP Vertica supports a series of built-in analytics libraries like time series and analytics packs for geospatial and sentiment plus additional functions from vendors like
SAS. And, it supports analytics written using the R programming language for predictive modeling.
HP Vertica is a Massively Parallel Processing (MPP) platform that distributes its workload over multiple commodity servers using a
shared-nothing architecture

Enables organizations to manage and analyze massive volumes of structured and semi-structured data quickly and reliably with no limits or business compromises.

What About Hadoop?
The HP Vertica Analytics Platform and Hadoop are highly complementary systems for big data analytics. The HP Vertica Analytics Platform is ideal for interactive, blazing-fast analytics and Hadoop is well suited for batch-oriented data processing and low-cost data storage.
When used together, with tight integration and support for all Hadoop distributions, the HP Vertica Analytics Platform and Hadoop offer the most open SQL on Hadoop. As a result, organizations can use the most powerful set of data analytics capabilities and do far more than either platform could do on its own, extracting significantly higher levels of value from massive amounts of structured, unstructured, and semi-structured data.

BM Netezza's approach to data warehousing
IBM's approach to data warehousing is very different from what is offered by vendors such as Oracle and Teradata. The IBM Netezza data warehouse appliances are purpose-built for crunching massive volumes of data quickly and efficiently. This enables organizations to realize business value quickly, and analytically explore areas previously unimagined. Competitive offerings just aren't able to do that.
Those traditional vendors are now offering their own versions of systems that are similar to IBM Netezza data warehouse appliances. For the most part, however, these are simply a repackaging of current technology and come with the same limitations: poor performance, complexity, high administrative costs and lack of scale.

In addition, every IBM Netezza data warehouse appliance is delivered with IBM Netezza Analytics, an embedded software platform for advanced analytics. It provides the technology infrastructure to support enterprise deployments of parallel, in-database analytics. Support for a variety of popular tools and languages as well as a built-in library of parallelized analytic functions make it simple to move analytic modeling and scoring inside the data warehouse appliance. IBM Netezza Analytics is fully integrated into the IBM Netezza data warehouse asymmetric massively parallel processing (AMPP) architecture enabling data exploration, model-building, model-diagnostics and scoring with unprecedented speed.

Oracle and RDBMS database
Its row and column based database used only of storing of highly structured database. Applying analytic and Processing on large amount of data for analytic purpose is very time consuming based on its design(row and column). To overcome from this situation, we started building cubes (OLAP) but this does not help us to do real time analytics.

Over the period Oracle the owner of Oracle RDBMS database realized the power of Big data and came up with Oracle Big data Appliance 
Oracle Big Data Appliance is an engineered system that combines optimized hardware with a
comprehensive big data software stack to deliver a complete, easy-to-deploy solution for
acquiring and organizing big data.
Oracle Big Data Appliance comes in a full rack configuration with 18 Sun servers for a total
storage capacity of 648TB. Every server in the rack has 2 CPUs, each with 8 cores for a total of
288 cores per full rack. Each server has 64GB1 memory for a total of 1152GB of memory per

full rack.
The Oracle Big Data Appliance software includes:
- Full distribution of Cloudera’s Distribution including Apache Hadoop (CDH4)
- Oracle Big Data Appliance Plug-In for Enterprise Manager
- Cloudera Manager to administer all aspects of Cloudera CDH
- Oracle distribution of the statistical package R
- Oracle NoSQL Database Community Edition2

- And Oracle Enterprise Linux operating system and Oracle Java VM

Reference:
http://www.oracle.com/us/products/database/big-data-for-enterprise-519135.pdf
http://www.netezza1000.com/

http://www8.hp.com/h20195/V2/GetPDF.aspx/4AA5-8917ENW.pdf