Function
Cloud & Data EngineeringOur Company
We’re Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world’s potential. We’re people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what’s now to what’s next. We make it happen through the power of acceleration.
Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don’t expect you to ‘fit’ every requirement – your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us.
Job description
Job Summary
We are seeking a highly skilled Hands-On Data Engineering Lead with deep expertise in Apache Flink, AWS Managed Service for Apache Flink, Kafka, and AWS. This role requires a technical leader who actively architects, develops, debugs, optimizes, and supports production-grade real-time streaming platforms. The ideal candidate combines hands-on engineering depth with experience leading teams that deliver scalable, resilient, high-throughput, and low-latency data solutions in 24x7 environments.
Key Responsibilities
Hands-On Technical Leadership
- Lead and mentor data engineers and architects while remaining directly involved in solution design, implementation, and troubleshooting.
- Serve as the technical authority for Apache Flink, AWS Managed Service for Apache Flink, Kafka-based streaming, and associated AWS services.
- Lead architecture, design, code, configuration, deployment, and production-readiness reviews.
- Establish engineering standards, coding practices, test discipline, and production-support procedures.
- Coordinate technical decisions across application, data, cloud-platform, performance, and operations teams.
Real-Time Streaming Architecture & Engineering
- Architect, design, and develop production-scale streaming solutions using Apache Flink, AWS Managed Service for Apache Flink, Java or PyFlink, Flink SQL, and Kafka.
- Apply stateful and event-time processing patterns, including keyed and broadcast state, state TTL, timers, windows, joins, watermarks, and delivery-semantics controls.
- Design resilient checkpointing, savepoint, restart, recovery, and state-migration strategies.
- Design Kafka topics, partitions, shard keys, routing, consumer groups, offsets, transactions, schemas, and source/sink integrations.
- Design and operate sharded streaming architectures, including workload decomposition, shard-key selection, state distribution, rebalancing, cross-shard processing, parallelism, failure isolation, and recovery.
- Build fault-tolerant streaming applications that meet defined throughput, latency, availability, and recovery objectives.
Performance Engineering & Optimization
- Define measurable entry, exit, and acceptance criteria for load, peak, burst, replay, recovery, and long-running soak tests.
- Verify that test inputs, replay behavior, duration, measurement windows, metric units, and outputs are representative and comparable.
- Analyze throughput, latency, backpressure, state growth, checkpoints, recovery, resource utilization, data skew, and operating headroom.
- Identify operator and stage-level bottlenecks and implement validated code, configuration, partitioning, or scaling improvements.
- Document test conditions, findings, qualifications, risks, and recommendations using reproducible evidence.
Production Engineering & Troubleshooting
- Act as a senior escalation point for complex production issues involving Apache Flink, AWS Managed Service for Apache Flink, and Kafka.
- Diagnose checkpoint failures, savepoint recovery issues, backpressure, state growth, idle partitions, data skew, memory pressure, garbage collection, restarts, and throughput or latency degradation.
- Analyze JobManager and TaskManager events, runtime configuration, logs, metrics, deployment history, and service behavior.
- Lead evidence-based root-cause analysis and define corrective and preventive actions for material incidents.
- Use controlled experiments to confirm or reject technical hypotheses and validate remediation effectiveness.
- Engage AWS Support and service specialists when deeper platform analysis or service-limit clarification is required.
Observability & Operational Readiness
- Define and implement monitoring for throughput, lag, backpressure, state size, checkpoints, restarts, failures, service events, and recovery.
- Standardize metric definitions, units, aggregation windows, data sources, thresholds, alerting, and escalation paths.
- Develop and review deployment, incident-triage, rollback, snapshot/savepoint recovery, scaling, and change-control procedures.
- Prepare technical documentation, operational runbooks, and knowledge-transfer materials for engineering and support teams.
Software Engineering Excellence
- Write, debug, test, optimize, and deploy production-grade Java, Python/PyFlink, and SQL code.
- Perform code reviews and promote automated testing, code quality, infrastructure-as-code, and DevOps practices.
- Build and improve CI/CD pipelines and deployment automation for streaming applications.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent professional experience.
- 10+ years of software engineering, data engineering, or distributed-systems experience, including 7+ years of recent hands-on Apache Flink engineering.
- Deep production experience with AWS Managed Service for Apache Flink, including deployments, runtime behavior, scaling, service limits, snapshots, monitoring, troubleshooting, and recovery.
- Strong command of the Flink DataStream API and Flink SQL, including stateful processing, event time, watermarks, windows, joins, timers, checkpoints, savepoints, and exactly-once concepts.
- Demonstrated experience diagnosing backpressure, checkpoint behavior, state growth, skew, idle partitions, memory and GC issues, restarts, Kafka lag, and performance degradation.
- Strong programming skills in Java, Python/PyFlink, and SQL, with experience developing and supporting production-grade streaming applications.
- Hands-on experience with Kafka or Confluent Kafka and sharded distributed systems, including topic and shard-key design, partitioning, routing, consumer groups, rebalancing, state migration, cross-shard processing, failure isolation, and recovery.
- Working knowledge of AWS services and controls supporting streaming platforms, including CloudWatch, S3, IAM, networking, service quotas, CI/CD, and infrastructure as code.
- Experience building and operating highly available, high-throughput, low-latency platforms in 24x7 production environments.
- Experience with monitoring, alerting, observability, incident response, root-cause analysis, and recovery planning.
- Ability to communicate technical findings clearly and distinguish observed facts, estimates, hypotheses, proposals, and approved decisions.
- Demonstrated ability to lead technical teams while remaining hands-on in engineering and troubleshooting activities.
Preferred Qualifications
- Experience supporting global, mission-critical streaming platforms and mentoring engineering teams.
- AWS certification or demonstrated equivalent AWS platform expertise.
- Experience processing very large event volumes using multi-shard architectures and managing multi-terabyte state, workload skew, state migration, or disaster-recovery design.
About us
We’re a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you.
Fostering innovation through diverse perspectives
Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth.
We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work.
How we look after you
We help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We’re also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We’re always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you’ll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with.
We’re proud to say we’re an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
Learn more about this Employer on their Career Site
