Company Overview:
We are building Protege to solve the biggest unmet need in AI — getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data, starting in the healthcare industry.
Solving AI's data problem is a generational opportunity. The company that succeeds will be one of the largest in AI — and in tech.
We're hiring for multiple data engineers ranging from senior to staff level to join the team.
Key Responsibilities and Scope
-
Work with the engineering team to design and implement scalable, automatable and robust data orchestration and lineage processes, incorporating high standards of data validation, transformation, and compliance.
-
Optimize data storage and retrieval costs, making data operations more efficient and cost-effective.
-
Work as a generalist engineer to help build other areas of the Protege platform.
About You
-
You are curious, tenacious, and proactive.
-
You are not bothered by ambiguity but embrace finding patterns in complex environments.
-
Proven experience in data engineering, with a solid track record of building data systems and processes.
-
Strong technical proficiency with data engineering tools, data warehouses and data lakes.
-
Strong programming skills in Java or Python.
-
Excellent problem-solving skills and adaptability in a dynamic and evolving tech landscape.
-
Excited to work in a company that deals with moving and transforming large volumes of data.
Bonus if you have these attributes
-
Experience with either Snowflake or Databricks
-
Experience with AWS
-
Prior startup experience
-
Familiarity with fine-tuning AI models
Top Skills
Similar Jobs
What you need to know about the Seattle Tech Scene
Key Facts About Seattle Tech
- Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Amazon, Microsoft, Meta, Google
- Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
- Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Madrona, Fuse, Tola, Maveron
- Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute