Pika Logo

Pika

Data Engineer

Posted 5 Days Ago
Remote
Hiring Remotely in US
Mid level
Remote
Hiring Remotely in US
Mid level
Design, build, and scale data pipelines, ETL workflows, analytics infrastructure, and machine-learning data systems. Ensure data quality, security, monitoring, and reliability while optimizing storage and processing performance. Collaborate with engineering, product, and analytics teams on data requirements, modeling, schemas, and scalable infrastructure. The role also establishes data engineering best practices and supports a data-driven culture.
The summary above was generated by AI

About Pika

 

At Pika, we’re building the next generation of AI creative tools to empower human creativity. Our mission is to make video creation seamless, intuitive, and accessible to everyone, leveraging the power of advanced AI. We believe that AI should amplify creative expression—enabling everyone to create, collaborate, and communicate across media. Our team includes engineers, artists, and product thinkers, all passionate about building tools that unlock new creative possibilities.

 

Pika has raised significant funding and is backed by leading investors, with a collaborative culture based in Palo Alto, CA. We prefer hybrid in-office, sharing ideas and launching products together.

 

About the Role

 

We are seeking a Data Engineer to design, build, and scale the data infrastructure powering Pika’s creative AI platform. As a Data Engineer, you will play a key role in architecting, implementing, and maintaining our data pipelines and analytics systems, enabling our team to make data-driven decisions and deliver world-class AI experiences. You will work closely with product, engineering, and data teams to ensure data is accurate, reliable, and accessible for users and internal business needs.

 

You will combine software engineering know-how with data architecture expertise, helping us build robust, scalable, and high-performance systems. Your contributions will directly support the success of millions of creators and help shape the future of AI-powered media tools.

 

What You’ll Do

 
  • Design, develop, and maintain scalable data pipelines and ETL workflows

  • Build, automate, and optimize our data infrastructure for analytics, reporting, and machine learning applications

  • Ensure data quality, consistency, and security across all sources and sinks

  • Collaborate with engineering, analytics, and product teams to define data requirements and deliver reliable datasets

  • Implement monitoring solutions and proactively resolve data pipeline issues

  • Optimize storage and data processing performance for growth and efficiency

  • Contribute to data modeling efforts and schema design for analytics and product needs

  • Help establish best practices and empower a data-driven culture across the organization

 

What We’re Looking For

 
  • 4+ years of experience as a data engineer or in a similar role designing, building, and maintaining data infrastructure

  • Strong software engineering background with proficiency in Python, SQL, and/or similar languages

  • Hands-on experience with data pipeline orchestration tools (Airflow, Prefect, Dagster, etc.)

  • Experience with cloud data platforms (AWS/GCP, Redshift, BigQuery, Snowflake, etc.)

  • Knowledge of database systems, data modeling, and data warehousing best practices

  • Familiarity with monitoring, logging, and data quality practices for data workflows

  • Excellent analytical and problem-solving skills with attention to detail

  • Great communication skills and ability to work cross-functionally in a collaborative environment

  • Self-motivated, curious, and comfortable in a fast-paced, high-growth startup

 

Nice to Have

 
  • Experience supporting data for machine learning or AI-powered applications

  • Familiarity with real-time or streaming data architectures (Kafka, Kinesis, etc.)

  • Prior work at high-growth startups or experience with rapid scaling

  • Open source, hackathon, or data engineering community experience

 

Our Stack

 

Python, Go, Node.js, Postgres, Redis, Docker, Kubernetes, AWS/GCP

 

What We Offer

 
  • Competitive salary in the AI industry

  • Substantial equity in a fast-growing startup defining the future of AI and creativity

  • Comprehensive health benefits, monthly stipends, and company retreats

  • Collaborative, high-growth culture—everyone contributes to growth and success

Similar Jobs

Yesterday
In-Office or Remote
73K-130K Annually
Junior
73K-130K Annually
Junior
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Design, develop, test, and maintain large-scale healthcare data pipelines and AWS data warehouse solutions. Build Redshift and EMR ETL processes, monitor pipeline performance, resolve data discrepancies, support deployments, conduct code reviews, and document database designs and testing. Collaborate with engineers, analysts, consultants, and cross-functional teams to improve data quality, scalability, and processing efficiency.
Top Skills: Amazon EmrAmazon RedshiftAmazon S3Automated Testing FrameworksAWSCi/CdETLOraclePythonSQLSQL Server
2 Days Ago
Remote or Hybrid
OH, USA
Senior level
Senior level
Financial Services
Build and operate scalable Databricks-on-AWS data pipelines using PySpark, Delta Lake, and lakehouse patterns. Optimize performance, implement data quality, monitoring, alerting, and automated remediation, and deliver curated datasets for BI and analytics partners. Collaborate with stakeholders on architecture and design while applying secure software engineering, CI/CD, agile, and operational stability practices. The role also uses AI-assisted development tools and supports workforce data analytics.
Top Skills: AlteryxAmazon AthenaAmazon EmrAmazon S3Apache AirflowApache IcebergSparkAutosysAWSAws CloudwatchAws GlueAws LambdaBitbucketClaudeDatabricksDatabricks WorkflowsDelta LakeDelta Live TablesGitGithub CopilotJavaJenkinsOracleParquetPysparkPythonScalaSigmaSpinnakerSQLTableau
2 Days Ago
Remote
United States
Mid level
Mid level
Artificial Intelligence • Information Technology • Professional Services • Software • Analytics • Generative AI • Big Data Analytics
Lead enterprise-scale Akeneo PIM implementations, including solution architecture, platform configuration, data architecture, API development, and third-party integrations. Collaborate with developers, architects, clients, and stakeholders to deliver digital transformation projects. Ensure data privacy, security, and lifecycle protection while documenting technical solutions, gathering requirements, managing expectations, and mentoring junior team members.
Top Skills: Adobe CommerceAkeneo PimAmazon DynamodbAws AppsyncAws LambdaCcpaDamGdprGraphQLMagentoMdmPimRest ApisShopify

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account