
Member of Technical Staff, Frontier AI | $100-$130/hr Remote
Overview
A remote senior contributor role focused on bridging research, data, and live AI systems. You’ll own evaluation, failure analysis, and iterative improvements to boost model and system performance. Expect close collaboration with researchers and operators to turn experimental signals into real-world gains and measurable impact. This position emphasizes hands-on ownership, rigorous validation, and clear communication of results across technical and non-technical stakeholders.
What You'll Do4
- 1Lead end-to-end research and evaluation cycles, from framing problems to validating signals and calibrating data quality.
- 2Design ML oriented data platforms, including task definitions, annotation schemas, rubrics, incentives, and data pipelines aimed at improving downstream model outcomes.
- 3Investigate model and system failures to uncover root causes, edge cases, and actionable improvement opportunities.
- 4Convert ambiguous real-world behavior into structured evaluation frameworks and new data categories, ensuring robust measurement.
Requirements5
- 1Strong judgment about when research signal is ready for external communication and publication
- 2Experience building ML datasets, evaluation frameworks, and quality assurance processes
- 3Ability to translate messy real-world behavior into concrete research questions and evaluation plans
- 4Comfort operating in ambiguity with ownership mindset and decisive action
- 5Excellent written and verbal communication, especially around tradeoffs, limitations, and signal strength
Who Should Apply
The ideal candidate is a technically seasoned problem solver who thrives at the intersection of research and production systems. You should excel at designing data-centric evaluation methods, communicating complex ideas clearly, and driving cross-functional work with researchers, domain experts, and operators to deliver defensible, impactful improvements.
Salary Insight
Base salary range discussed for similar roles is $180,000 to $320,000, with equity and performance-based bonuses often available depending on the role and policy.
Location
Required Skills
Application Tip
Prepare a concise portfolio of 2–3 projects where you shaped data definitions, evaluation metrics, and a deployment-ready improvement with measurable impact, and be ready to explain tradeoffs and data quality decisions.
Similar open positions
Explore active roles that match your skills and interests.

Micro1
VerifiedMember of Technical Staff, Enterprise AI | $100-$130/hr Remote
A remote, full-time role focused on advancing enterprise AI systems through hands-on research embedded in real workflows. You’ll identify real-world failure modes, run rapid experiments, and translate findings into impactful improvements. Expect to design data and evaluation strategies, build practical tooling, and contribute to external research artifacts that drive robust, scalable AI solutions.

Micro1
VerifiedMember of Technical Staff, Research Engineering | $7-$8/hr Remote
A research driven role focused on advancing reinforcement learning systems from concept to production. You’ll build novel RL environments, scalable training pipelines, and automated evaluation tools to push model capabilities. The work blends research ideas with robust, high-performance systems, all remote and full-time.

Micro1
VerifiedAI Evaluation Analyst | $20-$30/hr Remote
The AI Evaluation Analyst helps advance frontier language models from a remote setting. You’ll craft detailed, task-based conversations and evaluation rubrics, test them against leading LLMs, and produce clear, evidence-backed assets that guide model learning. This contract role centers on written clarity, multi-turn design, and meticulous analysis to shape how AI systems reason and respond.

Micro1
VerifiedMember of Technical Staff, Forward Deployed (US Gov) | $40-$60/hr Remote
A forward deployed engineering role focused on applying AI for mission-critical government work. You’ll move ideas from prototype to production, building agentic systems, LLM apps, and scalable data pipelines in high assurance environments. The position blends hands-on development with direct collaboration with government partners to turn needs into capable tech solutions.
Chakra-Labs
VerifiedForward Deployed Engineer, AI Data & Evals
You will own the full lifecycle of frontier AI research deliverables: from a researcher's hunch to a shipped environment, eval, or dataset. You will build with TypeScript and Python across backend services, data pipelines, and a usable frontend. You join a small team of ex-Stripe, Snap, AWS, and Airtable engineers who ship to named customers at frontier labs. This role compounds one-off builds into the core platform.
CHAOS Industries
VerifiedStaff Software Engineer Applied AI, Defense Tech & Real-Time AI Systems
AI powered defense and critical infrastructure systems require rapid prototyping and reliable deployment. Lead the design, integration, and shipment of AI/ML capabilities across CHAOS product lines at scale. Collaborate with researchers, engineers, and customers in a high-trust, high-impact environment where independence drives real results. Bring a track record of delivering practical AI solutions in regulated or mission-driven settings and mentor junior engineers on the team.