logo

JobNob

Your Career. Our Passion.

Site Reliability Engineer - DataPlatform


PhonePe


Location

Bangalore | India


Job description

Job Overview:

As a Site Reliability Engineer (SRE) specializing in Data Platform OnPremise, you will play a critical role in deployment, ensuring the reliability, scalability, and performance of our

Cloudera

Data Platform (CDP) infrastructure. You will collaborate closely with cross-functional teams to design, implement, and maintain robust systems that support our data-driven initiatives. The ideal candidate will have a deep understanding of Cloudera Data Platform, strong troubleshooting skills, and a proactive mindset towards automation and optimization. You will play a pivotal role in ensuring the smooth functioning, operation, performance and security of large high density Cloudera-based infrastructure.

Key Responsibilities:

Implementation of Cloudera Data Platform: Lead the implementation process of Cloudera Data Platform on-premises, including planning, installation, configuration, and integration with existing systems. Infrastructure Management: Manage and maintain the Cloudera-based infrastructure, ensuring optimal performance, high availability, and scalability. This includes monitoring system health, troubleshooting issues, and performing routine maintenance tasks. Data Security and Compliance: Implement and enforce security best practices to safeguard data integrity and confidentiality within the Cloudera environment. Ensure compliance with relevant regulations and standards (e.g., GDPR, HIPAA, DPR). Performance Optimization: Continuously optimize the Cloudera infrastructure to enhance performance, efficiency, and cost-effectiveness. Identify and resolve bottlenecks, tune configurations, and implement best practices for resource utilization. Capacity Planning: Monitor resource utilization trends and plan for future capacity needs. Proactively identify potential capacity constraints and propose solutions to address them. Backup and Disaster Recovery: Implement robust backup and disaster recovery strategies to ensure data protection and business continuity. Test and maintain backup and recovery procedures regularly. Patches & Upgrades: Routinely apply recommended patches and perform rolling upgrades of the platform in accordance with the advisory from Cloudera, InfoSec and Compliance. Documentation and Knowledge Sharing: Create comprehensive documentation for configurations, processes, and procedures related to the Cloudera Data Platform. Share knowledge and best practices with team members to foster continuous learning and improvement. Collaboration and Communication: Collaborate effectively with cross-functional teams including data engineers, developers, and IT operations personnel. Communicate project status, issues, and resolutions clearly and promptly.

Qualifications: Bachelor's degree in Computer Science, Engineering, or related field. Proficiency in Linux system administration, shell scripting, and networking concepts. 5+ years of experience in managing Big Data infrastructure. Strong understanding of distributed computing principles and experience with Hadoop ecosystem technologies (HDFS, MapReduce, YARN, Hive, Spark, etc.). Hands-on experience with configuration management tools (e.g., Salt,Ansible, Puppet, Chef). Strong scripting skills (e.g., Python, Bash) for automation and troubleshooting. Experience with monitoring and logging solutions (e.g., Prometheus, Grafana, ELK stack). Knowledge of networking principles and protocols (TCP/IP, UDP, DNS, DHCP, etc.). Experience with managing *nix based machines and strong working knowledge of quintessential Unix programs and tools (e.g. Ubuntu, Fedora, Redhat, etc.) Excellent communication skills and the ability to collaborate effectively with cross-functional teams. Excellent analytical, problem-solving, and troubleshooting skills.. Proven ability to work well under pressure and manage multiple priorities simultaneously.

Good To Have: Cloudera Certified Administrator (CCA) or Cloudera Certified Professional (CCP) certification preferred. Minimum 5 years of experience in managing and administering medium/large hadoop based environments (>100 machines), including Cloudera Data Platform (CDP) experience is highly desirable. Familiarity with Open Data Lake components such as Ozone, Iceberg, Spark, Flink, etc. Familiarity with containerization and orchestration technologies (e.g. Docker, Kubernetes, OpenShift) is a plus


Job tags



Salary

All rights reserved