Careers at Megahit

Build agents that own
the outcome.

Most agents are evaluated in seconds. Ours are judged three months later, by whether a company grew. Come solve the problems that gap creates.

Engineering Hiring now

AI Engineer, Agent Systems

Most agents are evaluated in seconds. Ours are judged three months later, by whether a company grew. Come solve the problems that gap creates.

Apply now →
Location
Remote, global
Type
Full-time
Team
Six people
Apply to
k@megahit.ai
The bet

More agent employees than human employees.

Every company is going to end up with more agent employees than human employees. Right now the human is the glue holding them together — a person sitting in front of a thousand open terminals, hand-carrying context between them. That doesn't scale, and nobody has built the layer that replaces it.

Coding agents were the easy case. The loop closes in seconds, the test either passes or it doesn't, and the whole industry has benchmarks. Every other kind of agent work — growth, finance, operations — has none of that. No business dataset, no evaluation set, no ground truth. That's exactly why adoption there is years behind, and exactly why the infrastructure for it is still unbuilt.

We're building it, starting with growth, because growth is the rare domain that is measurable but not uniform: every company has a different North Star, and the results take weeks or months to arrive. Hard enough that solving it is a real moat. Not so hard that it's a research project with no customer.

The actual work

Five problems we have not solved.

This is not a roadmap of tickets — it is a list of things that are genuinely open, that you would own, and that very few engineers anywhere have had to solve.

01

Context consistency

Twenty agents, two months, one coherent memory

A campaign runs for eight weeks across many agents. Some context is written by humans, some by agents, some is raw output, some is hard-won learning. By week six an agent is acting on something it concluded in week one — which may no longer be true.

What does an agent keep, what does it compress, what is it allowed to forget — and how does it know when a past conclusion has expired?

02

Judgment

The decision a planner cannot avoid

A sub-agent reports back: 30% behind goal, one week in. There are four real options — kill the channel and move the budget, change strategy and retry, renegotiate the goal with the human, or ask for more budget. A good human operator agonizes over this. There is no correct answer in the training data.

How do you build a planner that makes that call, explains why, and gets measurably better at it over time?

03

Delayed reward

Evaluating work whose result arrives in three months

SEO tells you whether you were right a quarter later. You cannot unit-test that, and you cannot wait for it either. The proxy signals have to be trustworthy enough to steer on, and honest enough to admit when they're wrong.

What is the benchmark for a task whose ground truth doesn't exist yet — and how do you prove a change made the system better rather than merely different?

04

Fleet drift

Sub-agents that learn separately without pulling apart

Each agent learns locally from its own channel. They report up, sync sideways, and adapt. Left alone, a fleet either diverges into contradictory strategies or averages itself into mush. Both failure modes look fine for the first two weeks.

How do you detect drift across a fleet before a human notices, and correct it without flattening what each agent actually learned?

05

Durability

Runs that survive contact with reality

Six hours in, hundreds of tool calls deep, mid-flight deploy, a third party down at 3am. The run has to resume rather than restart, and any past run has to be replayable step by step — otherwise nobody can say why the system did what it did.

What has to be journaled, and what does it cost, for an eight-week agent process to be fully explainable after the fact?

Why this is tractable here

Hard problems, grounded in real work.

Two reasons this isn't just a hard problem stated loudly.

First, we run the work ourselves for real customers before we automate it, which means our data records the decisions and the reasoning behind them — not just the outputs. Every agency on earth has static campaign data. Almost nobody has a record of why an operator chose to cut a channel instead of doubling down. That is the dataset the hard problems above actually need, and we've been building it from day one.

Second, we work with academic researchers on long-horizon context consistency and on the benchmarks that don't exist yet. If you want papers out of this as well as product, the door is open.

We have paying customers today. This is not a pre-product story, and it is not a research lab — the systems you build ship to people who are relying on them this week.
What the work looks like day to day

Own the system end to end.

Python, async-first, real concurrency. An in-house agent harness we own end to end and keep deliberately model-agnostic. Isolated sandboxes for agent-authored code. Journaled, replayable runs over a transactional store. Containerized services, deployed continuously. Connector protocols reaching dozens of external systems that break in creative ways.

Architecture specifics stay off the public internet. You'll get design docs, the trade-offs, and an honest list of what we got wrong in the first technical conversation — and you're expected to argue with all of it.

Who this is for

What will make you effective here.

Should be true

  • You've shipped an LLM agent real users depended on, and you can describe precisely how it failed.
  • Strong production Python: async, queues, and a database you've had to reason about under load.
  • You debug non-deterministic systems by instrumenting them, not by staring harder at the prompt.
  • You reach for an evaluation before you reach for an opinion — and you'll build the evaluation when none exists.
  • You're comfortable being the person who owns the hard part. There are six of us. There is nobody to escalate to.

Not required, but rare and valuable

  • You've built or seriously extended an agent framework rather than only used one.
  • Durable execution, workflow engines, or event-sourced systems.
  • Sandboxed or untrusted code execution.
  • Retrieval, memory, or context compression for long-running tasks — published or shipped.
  • Real curiosity about growth and go-to-market. Our engineers are honest that this is a gap, and closing it is worth a great deal here.

Who this is not for

  • You want a written spec. Here, half the job is deciding what the problem even is.
  • You want a mature codebase. Some of what you'll touch, you will end up replacing.
  • You want research without product pressure, or product without research. This role is both, in the same week.
Terms

Remote, global, and close to the decisions.

Remote and global, with a small distributed team across Asia and the US. No layer between you and the decision: you'll talk to customers, disagree with the founders, and see your work in production the same week you wrote it.

Competitive salary plus meaningful equity, sized for how early you're joining. We'll discuss numbers on the first call rather than making you guess.

Process

From first call to offer inside two weeks.

  1. 01

    A 30-minute call with KK, our CEO, on what you've built and what you want to build next.

  2. 02

    A technical conversation where we show you our architecture and you tell us what's wrong with it.

  3. 03

    A paid work sample, roughly a day, on a real problem from the list above.

  4. 04

    Offer. We aim to finish inside two weeks.

Apply

Pick a problem worth arguing about.

Send your resume, and pick one of the five problems above. In a paragraph or two, tell us how you'd attack it and what you think everyone gets wrong about it.

No cover letter. A repository, paper, or writeup beats a polished summary. A wrong answer argued well will get you further than a safe one. We read every application and reply either way.

Email k@megahit.ai →