Senior Data Engineer, Selling Partner Agentic Interfaces Data Products
Amazon
- Location
- IN, KA, Bengaluru
- Employment
- Full Time
- Work model
- On-Site
- Level
- Senior
- Posted
- Aug 28, 2026
Skills
About this role
Own the architecture of a data product that makes AI-powered commerce measurable, trustworthy, and improvable for millions of sellers worldwide. Join SP-AI Data Products as a founding technical leader who will define how an entire organization observes, governs, and learns from every AI-driven interaction at scale, built entirely on AWS large-scale data processing infrastructure. We're at an inflection point. AI agents are replacing traditional seller workflows, generating 7.5 million interactions annually across 50+ internal product teams, with thousands of Developers building on the ecosystem. Every one of these innovations creates a measurement obligation, and right now there's no unified infrastructure to fulfill it. You'll change that. You're joining a small team transforming from traditional reporting into a production data product organization, and this role determines what that product becomes. You'll design and build on AWS services including Kinesis for real-time streaming ingestion, EMR and Spark for large-scale distributed data processing, Glue for ETL orchestration, Redshift and Athena for analytical workloads, S3 and Lake Formation for governed storage, and Lambda and Step Functions for event-driven pipeline automation. What you'll own: - Design and govern the canonical event schema: the unified measurement format that makes every AI interaction across every surface produce a comparable, correlated record. You define what gets measured and how. - Own end-to-end data architecture across telemetry layers (user engagement, action execution, domain response), ensuring cross-layer correlation through session-level tracing - Build and operate streaming and batch ingestion pipelines on AWS with production SLAs, serving real-time observability for leadership, risk teams, applied scientists, and product teams simultaneously - Architect tenantized data products that enable 50+ teams to onboard once and receive self-serve metrics (adoption, quality, risk, impact) without building custom pipelines - Drive data engineering standards across the organization: schema governance, naming conventions, data quality, operational excellence. You set the bar others build to. - Make architectural trade-offs that balance short-term delivery against long-term scalability: tiered storage, build-vs-buy decisions, schema evolution across dozens of consumers - Identify and resolve systemic architecture deficiencies, proposing and leading cross-team initiatives that unblock innovation for adjacent teams - Decompose complex, ambiguous problems into parallel workstreams executable by you and others, then reassemble them into cohesive solutions - Elevate the engineering team through mentorship and technical leadership. Your presence makes the team stronger, but the team doesn't require your presence to succeed. Why you'll love this role: - Founding-team impact: You're defining the architecture an entire organization builds on, not inheriting legacy systems - Breadth of influence: Your schema decisions, quality standards, and architectural patterns are consumed by risk teams, scientists, product managers, and leadership across the business - Technical depth at scale: Streaming infrastructure, cross-surface correlation, schema governance for 50+ consumers, production SLAs on AWS. Hard, consequential engineering. - Career-defining scope: Cross-cutting schema design and multi-team domain onboarding at this scale is the kind of work that shapes what comes next If you've built large-scale data products on AWS that other teams depend on, thrive in ambiguity, and want to define how an organization measures AI-driven commerce at scale, we'd love to talk. Key job responsibilities - Own the design, implementation, and evolution of large-scale data architecture on AWS (Kinesis, EMR, Spark, Glue, Redshift, Athena, S3, Lake Formation), providing system-wide technical guidance and ensuring all data products meet production-grade reliability, scalability, and