| Senior Databricks Engineer Architect at Raleigh, North Carolina, USA |
| Email: [email protected] |
|
http://bit.ly/4ey8w48 https://jobs.nvoids.com/job_details.jsp?id=3360134&uid=fa046b57a31049f38c460151cb9b2e7c From: Mohsin, vyzeinc [email protected] Reply to: [email protected] Job Description - Senior Databricks Engineer Architect Years of Exp.: 13+ years exp. Work Location: Raleigh, NC/Local Remote - candidates must be local to Raleigh, NC. (For very senior candidates that meet all or more of the skills, clearances, etc. can be remote in the Eastern US.) Clearance: Public Trust Clearance or Higher Preferred Basic Qualifications: 5+ years of demonstrated experience designing and implementing data ingestion pipelines using tools such as Azure Data Factory, Apache Kafka, Apache NiFi, Spark Structured Streaming, or equivalent technologies. 5+ years of experience applying de-duplication techniques at scale, including record linkage, fuzzy matching, and entity resolution across structured and unstructured datasets. 5+ Hands-on experience with data tagging and metadata management, including the use of tagging schemas, data catalogs (e.g., Azure Purview, Apache Atlas), and automated classification tools to support data governance and lineage tracking. 5 + Demonstrated experience working with unstructured data. 2 + years of experience in using Databricks or other Spark-based platforms. Fluency in at least one scripting language (Python, Perl, Ruby, or equivalent). Integration of Git in continuous deployment and experience with DevOps monitoring tools. Experience with one or more of the following products and technologies: SAS, Python, C++, Hadoop, SQL Database/Coding, Teradata, Oracle, Amazon S3, Apache Spark, Machine Learning, Natural Language Processing, and visualization tools such as Tableau, Strategy and QLIK. Strong skills and experience in Cloud Operations support in Azure. Roles and Responsibilities (including but not limited to): Design, develop, and maintain scalable data ingestion pipelines to onboard structured, semi-structured, and unstructured data from batch and streaming sources (e.g., APIs, databases, flat files, message queues) into the Azure/Databricks environment. Implement de-duplication strategies across large-scale datasets using deterministic and probabilistic matching techniques to ensure data integrity and reduce redundancy within the Data Lake. Develop and enforce data tagging frameworks to classify, label, and annotate datasets with appropriate metadata (e.g., sensitivity, source, domain, lineage) to support data governance, discoverability, and compliance requirements. Assist with Operationalizing deployments and support of Cloud services for ETL Operations. This will include standardizing and automating processes and workflows, creating documentation/knowledge articles, and overall assisting Operations staff who have limited experience in Cloud. Written and oral presentations to high-level CIO management on status of current efforts. Possesses skills and experience related to business management, systems engineering, operations research, and management engineering. Typically has specialization in a particular technology or business application. Keeps abreast of technological developments and industry trends. Assist with deployment, configuration, and management of Azure Cloud environment. Assist with migration efforts of existing ETL jobs into Azure/Databricks cloud environment. Ability to share optimization and efficiencies with the larger team and management. Ability to automate solutions to repetitive problems/tasks. Keywords: cplusplus sthree information technology Delaware North Carolina Senior Databricks Engineer Architect [email protected] http://bit.ly/4ey8w48 https://jobs.nvoids.com/job_details.jsp?id=3360134&uid=fa046b57a31049f38c460151cb9b2e7c |
| [email protected] View All |
| 09:46 PM 08-May-26 |