
Senior Software Engineer, Data Platform (Remote-ok)
Tweets, press, and people that show why this team is worth your time.
Andrenam is hiring the first dedicated engineer for its data platform. Today the platform is nascent — solid backend infrastructure exists, but nobody owns it, and the perception/ML team is being handed data that isn't in the shape they need. This is a founding-level, zero-to-one seat: you'll define how raw maritime signals become clean, consistent, labeled datasets for perception and foundation models.
You'll architect high-throughput pipelines that ingest real-time acoustic and telemetry data, then align, resample, calibrate, and restructure it — not necessarily in the moment (it can be matched up every 30 minutes, hour, or day) but into exactly what the ML team needs. Expect to build backfill/reprocessing frameworks, lineage and versioning for reproducibility, dataset discovery APIs, data-quality instrumentation, and — likely — a labeling tool from scratch. There's real backend infra to build on, but the shape of the platform is yours to define.
What You'll Own
- Post-processing pipelines that align, resample, and calibrate multi-sensor data
- Backfill/reprocessing frameworks for new filters, syncs, label corrections, and metadata enrichment across historical data
- Lineage and versioning to guarantee experiment reproducibility
- Dataset discovery + access APIs/SDKs (query by time, region, modality, labels, quality flags)
- Data-quality metrics, dashboards, and alerts; canary dataset builds Storage-layout optimization (columnar formats, compression, chunking, sharding, prefetching)
- Likely build a labeling tool and pre-labeling workflows for the perception team
- Ramp: month 2 — a first working pipeline in place; months 3–6 — iterating and honing it to exactly what the ML/perception team needs, plus the tooling around it
Requirements
- Strong data-pipeline architecture, high-throughput, with the ability to architect the system from a blank page
- Solid knowledge of at least one cloud provider, preferably AWS
- Comfortable deploying pipelines to the cloud
- Python + data tooling (PyArrow/Polars/Pandas, NumPy/SciPy), plus one of Go/Rust/TypeScript for services
- Bonus: full-stack/generalist range — able to build the tools the ML team needs end-to-end
Execution and Ownership
- Comes in and builds day one with minimal hand-holding
- Low ego; takes criticism without taking it personally
- High autonomy; startup-native
Background
- ~5–8 years; startup time counts double, 7+ if the experience is at slower-moving shops
- Architected data pipelines / greenfield data-platform work; dataset-as-a-product ownership is a strong signal
- Exposure to edge/sensor data (audio/sonar, video, telemetry), time sync, and geospatial context is a plus
Location and Visa
- Remote OK; LA strongly preferred, with occasional on-site visits to Torrance
- US citizenship required
Nice-to-Have
- Labeling workflows (interfaces, ontologies, consensus, QA) and label-store integrations
- Splitting/sampling strategy design (by time, platform, geography, class, SNR) to avoid leakage
- Orchestration (Airflow, Prefect) and metadata/lineage (MLflow, W&B)
- Dataset-as-a-product track record with strong lineage and documentation
- Athletic or competitive background (e.g., competitive chess) — reads as dynamic and startup-fit
Who Will Thrive Here
- The data engineer who wants to own an entire platform employee-early
- Architect-operators who can go from block diagram to shipped pipeline without hand-holding
- Zero-to-one builders who've stood up data platforms from scratch and like greenfield
- Low-ego, autonomous, comfortable in crunch — "get your stuff done"
- Excited by the mission and the messy, real sensor data of the ocean