
Senior Software Engineer, Data Platform (Remote-ok)
Has experience building complex pipelines and storage systems within established environments, but lacks experience as the primary architect responsible for creating a foundational data platform from the ground up.
Experience is centered on AI and supporting data pipelines rather than personally architecting the foundational data infrastructure from scratch.
Looking for a senior engineer who has personally architected and owned a foundational data platform from scratch.
Plenty of impressive technology keywords but this is fundamentally an AI/ML engineer with data-pipeline exposure, not a senior 0 to 1 data/platform architect.
Strong backend/founding engineer with solid systems experience and greenfield product builds, but resume does not show ownership of a dedicated 0→1 data platform as required. Most data work seems focused on supporting their own application rather than building a reusable platform.
Does not demonstrate the kind of direct 0 to 1 data platform ownership we're looking for. Shows strong ingestion and infrastructure work, but not unmistakable ownership of building the core platform itself.
Doesn't demonstrate ownership of the kind of large-scale data platform the founders are searching for. Given the latest calibration, they're specifically looking for someone who's architected ingestion, storage, and the core data platform itself.
Looking for someone with more direct 0-1 data platform experience. Also has switched roles more frequently than preferred.
Strong software engineer with solid Spark, SageMaker, and data pipeline experience, but background is centered on building and optimizing individual pipelines rather than owning a greenfield data platform. Based on the client's latest calibration, they're looking for someone who has architected ing…
Background is centered on event-driven financial systems rather than owning a greenfield data platform or ML data infrastructure.
Tweets, press, and people that show why this team is worth your time.
Andrenam is hiring the first dedicated engineer for its data platform. Today the platform is nascent — solid backend infrastructure exists, but nobody owns it, and the perception/ML team is being handed data that isn't in the shape they need. This is a founding-level, zero-to-one seat: you'll define how raw maritime signals become clean, consistent, labeled datasets for perception and foundation models.
You'll architect high-throughput pipelines that ingest real-time acoustic and telemetry data, then align, resample, calibrate, and restructure it — not necessarily in the moment (it can be matched up every 30 minutes, hour, or day) but into exactly what the ML team needs. Expect to build backfill/reprocessing frameworks, lineage and versioning for reproducibility, dataset discovery APIs, data-quality instrumentation, and — likely — a labeling tool from scratch. There's real backend infra to build on, but the shape of the platform is yours to define.
What You'll Own
- Post-processing pipelines that align, resample, and calibrate multi-sensor data
- Backfill/reprocessing frameworks for new filters, syncs, label corrections, and metadata enrichment across historical data
- Lineage and versioning to guarantee experiment reproducibility
- Dataset discovery + access APIs/SDKs (query by time, region, modality, labels, quality flags)
- Data-quality metrics, dashboards, and alerts; canary dataset builds Storage-layout optimization (columnar formats, compression, chunking, sharding, prefetching)
- Likely build a labeling tool and pre-labeling workflows for the perception team
- Ramp: month 2 — a first working pipeline in place; months 3–6 — iterating and honing it to exactly what the ML/perception team needs, plus the tooling around it
Requirements
- Strong data-pipeline architecture, high-throughput, with the ability to architect the system from a blank page
- Solid knowledge of at least one cloud provider, preferably AWS
- Comfortable deploying pipelines to the cloud
- Python + data tooling (PyArrow/Polars/Pandas, NumPy/SciPy), plus one of Go/Rust/TypeScript for services
- Bonus: full-stack/generalist range — able to build the tools the ML team needs end-to-end
Execution and Ownership
- Comes in and builds day one with minimal hand-holding
- Low ego; takes criticism without taking it personally
- High autonomy; startup-native
Background
- ~5–8 years; startup time counts double, 7+ if the experience is at slower-moving shops
- Architected data pipelines / greenfield data-platform work; dataset-as-a-product ownership is a strong signal
- Exposure to edge/sensor data (audio/sonar, video, telemetry), time sync, and geospatial context is a plus
Location and Visa
- Remote OK; LA strongly preferred, with occasional on-site visits to Torrance
- US citizenship required
Nice-to-Have
- Labeling workflows (interfaces, ontologies, consensus, QA) and label-store integrations
- Splitting/sampling strategy design (by time, platform, geography, class, SNR) to avoid leakage
- Orchestration (Airflow, Prefect) and metadata/lineage (MLflow, W&B)
- Dataset-as-a-product track record with strong lineage and documentation
- Athletic or competitive background (e.g., competitive chess) — reads as dynamic and startup-fit
Who Will Thrive Here
- The data engineer who wants to own an entire platform employee-early
- Architect-operators who can go from block diagram to shipped pipeline without hand-holding
- Zero-to-one builders who've stood up data platforms from scratch and like greenfield
- Low-ego, autonomous, comfortable in crunch — "get your stuff done"
- Excited by the mission and the messy, real sensor data of the ocean