Software Engineer, Lakehouse (Data Platform Group)
Cato Networks
- Location
- Tel Aviv District, Israel
- Work model
- On-Site
- Level
- Mid
- Posted
- 2h ago
Skills
About this role
Welcome to the future of cloud networking and security!
Cato Networks is the first company to converge enterprise networking and security into one centralized and global service that is delivered by cloud. It is led by networking and security pioneer Shlomo Kramer (Check Point, Imperva) and early investor (Palo Alto Networks, Exabeam, Trusteer and more). Cato’s unique technology inspired a brand-new product category, later named “SASE” by Gartner and a market expected to reach $28.5 billion by 2028.
This is your opportunity to get on the rocket ship and join a company that is building a cutting-edge enterprise network and secure cloud platform, and is on a fast track to becoming the worldwide market leader – don’t miss it!
We're looking for an experienced Software Engineer to join our Data Platform Group. In this key role, you will build the company data platform: cloud-based microservices and data pipelines that process on the order of 1M records/sec at low latency. Because the platform is the foundation other groups build on, your work has a direct impact on our customers and enables engineering, product, and research teams across the organization. The Lakehouse team owns the data itself — how it is stored, organized, retained, and served. We own our analytical storage layer, the batch processing built on top of it, and the APIs through which customers and the rest of the company consume data. If you enjoy the problems that only appear at petabyte scale — physical data layout, query performance, storage cost, and retention — this is the role.
Responsibilities
• End-to-end ownership of our large-scale analytical storage layer: data modeling, schema and table design, partitioning, retention, and query performance.
• Design and develop the batch processing layer over our data lake using Spark on EMR (Java and PySpark).
• Build and evolve Java/Spring Boot services that expose our data through well-defined APIs to customers and to consumers across the company.
• Own performance and cost: query optimization, file layout and compaction, cluster sizing, and storage efficiency at scale.
• Research new technologies in the lakehouse and analytical-storage space and adapt them for use in our product.
• Work closely with product, DevOps, and security teams.
Requirements
• 5+ years of hands-on experience designing and developing large-scale distributed data systems in production, with a strong emphasis on performance.
• Deep, hands-on expertise in at least one of the following, at a significant scale:
• A columnar/analytical database — ClickHouse is a major advantage, including data modeling, query optimization, and operating it in production
• Apache Spark at an expert level, including tuning and optimizing large batch jobs.
• Experience with open table formats such as Iceberg, Delta Lake, or Hudi
• Experience with data lake technologies: Parquet, S3, and SQL query engines such as Athena, Trino, or Presto.
• Strong command of analytical data modeling and the design principles behind it: partitioning strategies, denormalization, batch vs. streaming trade-offs, and schema evolution.
• Strong Java and solid understanding of object-oriented design and software engineering principles.
• Experience building and running microservices on Kubernetes.
• Hands-on experience with the AWS platform, particularly EMR, S3, and Glue.
• Motivated, fast, independent learner and strong problem solver.
• A team player with excellent collaboration and communication skills.
• B.Sc. in Computer Science, Software Engineering, or a related field, or