| KARTHEEK REDDY - DATA ENGINEER |
| [email protected] |
| Location: Dallas, Texas, USA |
| Relocation: yes |
| Visa: H1b |
| Resume file: Kartheek Reddy Anugu-BDH-Resume_1788877170600.docx Please check the file(s) for viruses. Files are checked manually and then made available for download. |
|
SR. DATA ENGINEER
HADOOP | AWS | PYTHON | SCALA | SPARK | HIVE | ORACLE | SQL | LINUX PROFESSIONAL SUMMARY: 9+ years of experience in Development, Design, Integration, and Presentation with Java along with Extensive years of Big Data /Hadoop experience in Hadoop ecosystem such as Hive, Flume, Sqoop, Zookeeper, HBase, SPARK, Kafka, AWS, Map Reduce, Spark, Impala, Pig, Oozie, Scala, and Python. Experience in working with Data Frames, RDD, Spark SQL, Spark Streaming, APIs, System Architecture, and Infrastructure Planning. Extensive experience in importing and exporting data using stream processing platforms like Flume and Kafka. Expertise also in NoSQL databases like MongoDB, Map R-DB and Cassandra. Strong expertise on Amazon AWS EC2, Dynamo DB, S3, Kinesis and Amazon Athena, and Amazon Glue services Expertise in Big Data architecture like Hadoop (Azure, Hortonworks, Cloudera) distributed system, MongoDB, NoSQL Good experience in CI/CD pipeline management through Jenkins. Automation of manual tasks using Shell scripting. Good experience with Schedulers, like Oozie Workflows, CA workload, Autosys, and Airflow. Hands on databases like Oracle, MS SQL, DB2 and developing in RDBMS that includes SQL queries, Stored procedures and triggers. Strong knowledge in working with UNIX/LINUX environments, writing shell scripts and PL/SQL Stored Procedures. Experienced in using source code, control systems such as GIT, SVN, and CVS. TECHNICAL SKILLS: Big Data Technologies Hadoop, MapReduce, HDFS, Hive, Pig, Sqoop, Apache Spark, Scala, Yarn, and Apache Kafka, Azure Databricks Programming Languages C, C++, Java, Python, scala, SQL, PL/SQL. Frame works Struts, Spring, MVC. Cloud Platform AWS (EMR, EC2, S3), Microsoft Azure. Scripting Languages Shell Scripting, Java script, Linux, Unix. Application Build Tools Apache Ant, Apache Maven. Databases Oracle 9i/10g/11g, SQL Server, MySQL, Teradata, NoSQL Version Control System CVS, SVN, and GitHub. PROFESSIONAL EXPERIENCE: SR. DATA ENGINEER Public Consulting Group | DALLAS, TX September 2023 - Present Responsibilities: Setup Hadoop cluster Amazon EC2 ng whirr for POC. Worked on analyzing Hadoop Cluster and different big data analytic tools including Pig Hbase database and Sqoop. Responsible for building scalable distributed data solutions using Hadoop. Installed and configured Flume Sqoop Hbase Hive Pig on Hadoop cluster. Managing and scheduling jobs on Hadoop cluster. Implemented nine nodes CDH3 Hadoop cluster on Red Hat Linux. Worked on installing cluster commissioning decommissioning datanode namenode recovery capacity planning and slots configuration. Involved in loading data from UNIX file system to HDFS. Ingested data from RDBMS and performed data transformations, and then export the transformed data. to MongoDB as per the business requirement. Created Hbase tables to store variable data formats of PII data coming from different forms of portfolios. Installed and configured Hive and written HIVE UDFs. Cluster coordination services through zookeeper. Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team. Developed spark applications in Python (PySpark) on distributed environment to load huge number of CSV files with different schema in to Hive ORC tables. Extensively worked on Python and build the custom ingest framework. Implementing Spark using Python and Spark SQL for faster testing and processing of Data. Developed Spark applications using Scala & Python to do the analytics on the data stored on HDFS. Analyzed large amounts of datasets to determine optimal way to aggregate and report on it. Imported data using Sqoop to load data from MySQl to HDFS on regular basis. Implementing Spark using Python and Spark SQL for faster testing and processing of data. Handled Data-Skewness in Spark-SQL. SR. DATA ENGINEER BANK OF AMERICA | DALLAS, TX October 2022 July 2023 Responsibilities: Experience in Job management using Fair scheduler and Developed job processing scripts using Oozie workflow. Used Spark-Streaming APIs to perform necessary transformations and actions on the fly for building the common learner data model which gets the data from Kafka in near real time and Persists into Cassandra. Configured deployed and maintained multi-node Dev and Test Kafka Clusters. Developed Spark scripts by using Scala shell commands as per the requirement. Data pipelines for financial datasets in Azure Databricks. Delta Lake tables (ACID, schema enforcement, upserts). PySpark job optimization on Databricks clusters. Azure Data Factory + Databricks ETL orchestration across banking systems Experienced in performance tuning of Spark applications for setting right Batch Interval time, correct level of Parallelism and memory tuning. Database development required creation of new tables, PL/SQL stored procedures, functions, views, indexes and constraints, triggers and required SQL tuning to reduce the response time in the application Experienced in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and other during ingestion process itself. Optimizing of existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frames and Pair RDD's. Implemented ELK (Elastic Search, Log stash, Kibana) stack to collect and analyze the logs produced by the spark cluster. Analyzed the SQL scripts and designed the solution to implement using PySpark. Good experience with Talend open studio for designing ETL Jobs for Processing of data. implemented Partitioning, Dynamic Partitions, Buckets in HIVE. Good experience with continuous Integration of application using Jenkins. Used Reporting tools like Tableau to connect with Hive for generating daily reports of data. O Collaborated with the infrastructure, network, database, application and BI teams to ensure data quality and availability. DATA ENGINEER CITI BANK | DALLAS, TX December 2021- September 2022 Responsibilities: Writing complex Hadoop transformation jobs to extract the data which involves hive and impala transformations. Developed spark applications in Python (PySpark) on distributed environment to load huge number of CSV files with different schema in to Hive ORC tables. Writing Spark jobs to apply compaction on smaller HDFS block size files in HDFS for cluster consumption optimization. Building an internal Hadoop framework to load the data into Data Mart for Snapshot and Incremental tables along with aggregations and transformations. Developed fully customized framework using python, shell script, Sqoop & hive. Migrated data from Oracle to Data Lake using Sqoop, Spark and Talend (ETL Tool). Extensively worked on Python and build the custom ingest framework. Converted all Hadoop jobs to run in EMR by configuring the cluster according to the data size. Worked on Batch processing and Real-time data processing on Spark Streaming using Lambda architecture. Extensively worked on creating data pipeline integrating Kafka with spark streaming application used Scala for writing applications. Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data. Used Spark to process the data before ingesting the data into the HBase and both batch and real-time spark jobs were created using Scala. Developing and Managing Data Pipelines: Design, implement, and maintain complex data pipelines using Apache Airflow to ensure efficient and reliable data processing. Used SBT to build the Scala project. Working on CDC (Change Data Capture) tables using Spark Application to load data into Dynamic Partition Enabled Hive Tables. Documented logical, physical, relational, and dimensional data models. Designed the Data Marts in dimensional data modeling using star and snowflake schemas. Created Snow pipe for continuous data load and applied transformation logic using Snow SQL. Implementing effective optimization techniques to improve the query performance in Hadoop transformation queries. Schedule Hadoop jobs by writing Autosys scripts in Autosys automated job control system. DATA ENGINEER AMAZON | DALLAS, TX March 2021-November2021 Responsibilities: Setup and build AWS infrastructure various resources (AMI, VPC EC2, S3, IAM, EBS, Security Group, Auto Scaling, RDS, Glue, Glue data catalog, CloudTrail, Athena) using Packer and Terraform JSON templates. Working on AWS services (S3, EC2, ELB, EBS, Route53, VPC, Auto scaling etc.) and deployment services (Lambda, and Cloud Formation) and security practices (IAM, logging/CloudWatch and CloudTrail) and RDS, DynamoDB (NoSQL), Beanstalk, Cloud Front, ECS, SQS, SNS and SES. Working on version control system tool Bitbucket and GIT and having strong knowledge on source control concepts like Branches, Masters, Merges and Tags. Build Pipeline design and optimization: GIT, Maven, Nexus/Artifactory, Application servers and AWS for J2EE application deployments. Provides additional debugging support to the project team during development, production and enhancement work. Experience troubleshooting networking issues using several tools (traceroute, mtr, ping, iperf, dig/nslookup, cURL, tcpdump/wireshark and related) Knowledge of the Media Layer of the OSI Model (Physical, Data Link, and Network) components and troubleshooting. Knowledge of Networking (HTTP, SSL DNS, TCP/IP, IPSEC) and routing protocols (BGP). Hands on experience in setting up workflow using Apache Airflow and oozie workflow engine for scheduling and managing Hadoop jobs. Install and configure Apache Airflow for S3 bucket and snowflake data warehouse and create dags to run the Airflow. Automated resulting scripts workflow using Apache Airflow and shell scripting to ensure daily execution in production. Work on critical, highly complex problems that span multiple cloud computing services. Educate customers on the shared security model & assist them to secure their cloud environment. Advise customers on best practices of using AWS cloud technologies. Apply advanced troubleshooting techniques to provide unique solution to AWS customers needs. Troubleshoot cloud deployments, recreate customer issues and build proof of concept applications. Leverage day - to-day experiences to provide the voice of the customer to internal AWS teams. Maintain deep working knowledge of groundbreaking cloud computing technologies. Field and manage technical issues via phone, chat and email. Drive customer communication during critical events. Write tutorials, how-to videos and other technical articles for the AWS support community. DATA ENGINEER DUKE ENERGY | CHARLOTTE, NC October 2019-Febraury 2021 Responsibilities: Ingested data from RDBMS and performed data transformations, and then export the transformed data to MongoDB as per the business requirement. Wrote check code in Groovy to verify the correctness of map data. Analyzed failures found by the checks to determine exactly what was wrong. Assisted other developers with learning Groovy and advised them on writing efficient code. Migrated Hive QL queries on structured data into Spark QL to improve performance. Involved in improving the performance and optimization of the existing algorithms using Spark. Handled large datasets using Partitions, Spark in-memory capabilities, broadcasts in Spark, effective efficient Joins, transformations. Well experienced in handling Data-Skewness in Spark-SQL. Worked on the large-scale Hadoop Yarn cluster for distributed data processing and analysis using Spark, Hive, and HBase. Build code changes using CI/CD pipeline and deploy the kit on different environments. Database development required creation of new tables, PL/SQL stored procedures, functions, views, indexes and constraints, triggers and required SQL tuning to reduce the response time in the application. Building/orchestrating Spark ETL pipelines on Azure Databricks. PySpark/Scala notebooks integrated with ADLS Gen2. Cluster configuration and auto-scaling tuning. Databricks + Azure Data Factory orchestration. Implementing Spark using Python and Spark SQL for faster testing and processing of Data. Involved in creating data-lake by extracting customers data from various data sources to HDFS, which include data from Excel, databases, and log data from servers. Worked as a Spark Expert and performance Optimizer. Used Kafka functionalities like distribution, partition, replicated commit log service for messaging systems by maintaining feeds. Extensively worked on CAWA applications for running backend workflow of Hadoop environment. Used OOZIE Operational Services for batch processing and scheduling workflows dynamically. Used Zookeeper to coordinate the servers in clusters and to maintain the data consistency. Developed UNIX shell scripts to load large number of files into HDFS from local file system. HADOOP DEVELOPER BANK OF AMERICA | CHARLOTTE, NC July 2018 October 2019 Responsibilities: Involved in Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS using Java. Handled importing data from different data sources into HDFS using Sqoop and performing transformations using Hive, MapReduce and then loading data into HDFS. Captured the data logs from web server into HDFS using Flume for analysis Worked on Data Serialization formats for converting Complex objects into sequence bits by using AVRO, PARQUET, JSON, CSV formats. Managed Hadoop jobs by DAG using Oozie workflow scheduler. Involved in developing code to write canonical model JSON records from numerous input sources to Kafka Queues. Worked on structured and unstructured applications (SDAS & UDAS) to enhance the performance and eliminate the issues during development and testing. Working on current structured archiving project (SDAS) transition from PLAY app to complete Tomcat based application. Developed expandable rest client for the restful web services using Spring MVC, Java. Designed and developed the web-tier using Html, JSP's, Servlets and Tiles framework. Developed JSP for UI and Java classes for business logic. Developed a framework of RESTful web services using Spring MVC, JPA and APIs to help Hadoop developers to automate data quality checks. Designed HBase schema to avoid Hot spotting and exposed the data from HBase tables to REST API on UI. Data storage in HBase using Pig and Involved in Parsing of data using Pig. Involved in implementation of script to transform data from Oracle to HBase using Sqoop Responsible for performing extensive data validation between Hive tables and RDBS tables. HADOOP DEVELOPER CIGNA INSURNACE | BOSTON, MA January 2017 July 2018 Responsibilities: Migrated an existing on-premises application to AWS and used AWS services like EC2 and S3 for small data sets processing and storage, experienced in maintaining the Hadoop cluster on AWS EMR. Imported data from AWS S3 into Spark RDD, performed transformations and actions on RDD's. Worked on the large-scale Hadoop Yarn cluster for distributed data processing and analysis using Spark, Hive, and HBase. Involved in creating data-lake by extracting customer's data from various data sources to HDFS, which include data from Excel, databases, and log data from servers. Used Apache Solr to index the documents and used free-form queries to search the indexed documents. Developed Spark applications by using Scala and Python and implemented Apache Spark for data processing from various streaming sources. Developed Spark applications using Scala & Python to do the analytics on the data stored on HDFS. Worked as a Spark Expert and performance Optimizer. Written MapReduce (Hadoop) programs to convert text files into AVRO and loading into Hive (Hadoop) table. Worked with Spark for improving performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frames, and pair RDD's. Loaded D-Stream data into Spark RDD and did in-memory data computation to generate output response. Involved in loading data from rest endpoints to Kafka producers and transferring the data to Kafka brokers. Used Kafka functionalities like distribution, partition, replicated commit log service for messaging systems by maintaining feeds. EDUCATION: Master s in computer technology | 2017 Eastern Illinois University, Charleston, Illinois. Bachelor s in computer science engineering | 2015 Jawahar Lal Technological University, Hyderabad, India. Keywords: cprogramm cplusplus continuous integration continuous deployment user interface business intelligence sthree database rlang information technology microsoft mississippi procedural language California Massachusetts North Carolina Texas |