Posts

Spark - Optimal executors, cores and executor memory - Tuning spark jobs

Image
Executor, memory and core setting for optimal performance on Spark Spark is adopted by tech giants to bring intelligence to their applications. Predictive analysis and machine learning along with traditional data warehousing is using spark as the execution engine behind the scenes. I have been exploring spark since incubation and I have used spark core as an effective replacement for map reduce applications. Optimizing jobs in spark is a tricky area as there are no many ways to do it. I have done some trial and error in they way I write code sequencing. But as Spark use lazy evaluation and DAGs are pre created during execution there are no may ways to alter it.  This blog share some information about optimizing spark jobs - programmatically i.e. writing better code and playing around with hardware. Rule of thumb for better performance Use per key aggregation  - use reduce by instead of group by . Blog talks about the difference aggregate function options a...

Compose and Send HTML emails from Informatica (ETL)

Image
Dynamically generating and sending html emails using ETL(Informatica) Sending status emails from ETL is a very common practice in data warehouse projects. email tasks are available in all ETL tools which makes this task much easier. Normally, content of these emails are dynamic and created using Unix scripts or in some case ETL itself. Common ETL generated summary emails include Error reports ETA and job/load status Data warehouse/ Mart load completion Database/Server capacity alerts Summary reports in email Most of these emails are send to business group or IT project support itself. These emails are formatted and the real data is send as attachments(.csv,.xls ,.txt etc). Here, I am demonstrating a method to compose and send html emails using ETL and Unix command. I am deviating from usual method of sending emails with attachment and instead writing the attachment content into email body itself. Advanatge of html email is that , the data can be visually re...

Apache Spark -aggregate functions explained (reduceByKey, groupByKey and combineByKey)

Image
Easy explanation on difference between spark's aggregate functions (reduceByKey, groupByKey and combineByKey)   Spark comes with a lot of easy to use aggregate functions out of the box. For the same reason spark becomes a powerful technology for ETL on BigData. Grouping the data is a very common use case in world of  ETL(Extract , Transform and Load). Just like aggregate transformation in ETL tools like Ab-initio or Informatica, where the results can be grouped and aggregate functions can be applied. e.g. Group all customer order based on customer key, find the best sales year , find the worst player in baseball based on strike rate etc. etc. Unlike standard ETL tools sparks comes with three transformations to achieve the same result but in different ways. reduceByKey groupByKey combineByKey These are three transformation available in spark which can be used interchangeably. Before getting to further details, it is important to understand all this ...
Image
How to Insert data to remote Hive server from Spark Spark is the buzz word in world of BigData now. So what makes Spark so unique? As we know, Spark is fast - it use in memory computation on special data objects called RDD (Resilient distributed data set) Spark allows execution on multiple modes i.e. run standalone, run local (without even a hadoop server), on cluster through resource managers (Mesos, YARN) Spark take care of data lineage, fault recovery through DAG(Direct Acyclic Graph) as blue print for execution, which can be rebuilt at any point in case of failures Easy APIs - Easy to use APIs Read from anywhere - Data can be read from different types of sources i.e. files, json, databases etc. e.g. CassandraAPI Write to anywhere - Result data can be saved to any format Multiple language support - Spark supports scala, java , python. For people from database and SQL background Spark is simplified, to run SQLs on RDD (called dataframes) through SparkSQL(known as shark e...