David Komer · Founder & CEO
Test the candidate on the actual work
Jesse and the Gateway X team:
Lockstep is an assessment platform for realtime AI. Entire industries are still screening researchers and developers on static answers even though the job is to build a system that observes, decides, and acts while the world keeps moving.
With Lockstep, candidates submit a working binary or policy. They can use Rust, C++, Python, or another compatible toolchain. The submission runs as a sandboxed WASM component, with first-class ONNX inference when needed. Lockstep tests it across varied private scenarios under a fixed per-tick compute budget, then returns score distributions, failure modes, resource use, and replays the hiring team can inspect.
The core platform is live today. It combines sandboxed WASM, ONNX inference, 3D physics, per-tick metering, ladders, and replay in approximately 91,000 lines of Rust. The round funds the employer product on top: private campaigns, invitations, scenario suites, reports, candidate comparison, and data controls.
The starting plans use a budget employers already understand: $990 for 60 annual attempts and $4,500 for 300. From there, sponsored competitions and turnkey worlds grow the environment catalog. Production-agent evaluation becomes the larger second act.
I have spent more than twenty years building realtime, simulation, WebAssembly, rendering, and high-stakes financial systems, including eight years as a CTO. That background removed much of the platform risk before outside capital; I will personally own the first buyer conversations and paid campaigns.
The Gateway X Fellowship combines initial capital with twelve weeks working side by side in St. Louis. We would use that period to talk to customers, build the assessment layer around what we learn, package the first environments, and turn the live core technology into a company with repeatable demand.
Question 1 · What is it?
Evaluate the candidate on the system they build
Lockstep is a technical assessment platform for researchers and engineers who build realtime AI systems. A candidate's policy runs inside a changing world, so the employer can evaluate behavior—not another static coding answer.
Question 2 · Who is the customer—and what do we know?
Physical-AI teams can’t afford to confuse a good answer with a good hire
The first buyer is a technical leader hiring researchers and engineers for autonomy, controls, or policy roles—work where conventional screens can reward plausible code while missing robustness, recovery, efficiency, and realtime performance.
Series A–C physical-AI companies with 50–250 employees, multiple open autonomy or controls roles, and a Head of Autonomy, Robotics, Controls, or Engineering who owns the quality of the technical screen.
The need is visible
Multiple live roles, a recent deployment or funding event, and job descriptions calling for simulation-to-hardware transfer, recovery, latency, or realtime decision-making.
The stack is mixed
Researchers and engineers use Python to train and create policies; C++ and Rust engineers build controllers and deployment systems. Lockstep accepts both through one portable behavioral contract.
The decision is technical
The technical leader defines the bar, talent operations runs the campaign, and the hiring manager reviews replay-linked evidence instead of another generic score.
The most valuable hiring signal is not how a candidate explains an acting system. It is how the system actually acts.
BigID VP of Data Science & Analytics Yehoshua Enuka has tentatively offered product and implementation counsel grounded in AI/ML hiring needs. AI GTM specialist Moshe Porat has offered company introductions when the employer workflow is ready. David's direct access to the Farama community provides practitioner feedback and candidate-side distribution.
Question 3 · What is the white space?
The missing layer between coding tests and real-world simulation.
Employers can test coding skill, inspect notebook experiments, or build a company-specific simulator. None gives them a maintained assessment platform that accepts the mixed artifacts realtime teams actually build, runs them safely in live worlds, and turns behavior into comparable hiring evidence.
Python researchers can submit trained ONNX policies while Rust and C++ engineers submit native controllers. One portable WASM contract serves both sides of the realtime-AI stack.
Coding screens grade an answer after the fact. Lockstep exposes decisions, recovery, and stability as physics, opponents, and conditions change under a fixed realtime budget.
Lockstep ties robustness and failure modes to watchable replays, then shows per-tick resource use—including time spent in policy code versus inference.
Why this space remains open
Available tools solve one part. Employers must stitch the rest together.
Assessment platforms provide hiring workflow but not realtime system execution. Notebook tools evaluate models without deployment constraints. Internal simulators provide fidelity but consume scarce engineers to adapt, secure, calibrate, and maintain every exercise. Lockstep productizes the missing layer.
Untrusted code, safely
WASM provides a constrained capability set, fast startup, and near-native execution without giving every candidate artifact an operating system to attack.
Performance under pressure
The platform measures decisions as the world changes—not a final answer after the fact—including deadlines, recovery, and state over time.
Compute at the tick level
Because Lockstep owns the execution boundary, it can show exactly when a policy slowed down and whether time went to code or inference.
Failures you can watch
Engine and renderer are separate by design, turning any run into replay-backed evidence a hiring manager can inspect and discuss.
No technical gap stays empty. Lockstep's head start must become a catalog of calibrated environments, scenario patterns, campaign data, and employer trust. Each simulation can power public practice, a sponsored competition, private hiring suites, and eventually continuous agent evaluation.
Question 4 · Why now—and how big?
Coding tests are losing signal. Realtime AI roles are growing. The infrastructure is ready.
Static assessments are losing signal, autonomous and realtime systems are spreading across industries, and the infrastructure required to execute untrusted policies safely has become practical.
The required infrastructure arrived
WASI Preview 2, the WebAssembly component model, portable ONNX inference, and inexpensive sandboxing made live behavioral evaluation economical.
The static test broke
AI can generate conventional coding solutions, forcing employers to look for work samples that remain meaningful with AI assistance.
AI is moving into realtime systems
Robotics, autonomy, finance, games, and defense increasingly depend on systems judged by trajectories, recovery, latency, and consequences.
As AI leaves the chat window and enters the physical economy, the scarce skill becomes building systems that act well under time pressure and resource constraints.
2023–2030
2025–2033 estimate
and global members
These figures are not an additive TAM. The first two are third-party market estimates; Kaggle demonstrates the aggregate distribution available through competitions. Continuous agent evaluation remains the expansion thesis, supported by Arena's company-reported $100M annualized run rate and 5M+ monthly Agent Mode turns—not audited full-year revenue or a market-size estimate. Together they show the path from a multibillion-dollar hiring wedge into simulation, global participation, and recurring evaluation.
Question 5 · Show, don't tell
The core technology is live—and you can watch it run.
Dance Off demonstrates the foundation beneath every future assessment: compiled systems operating inside the same live physics world while Lockstep records both what they do and what they cost to run.

Submit
A WASM component compiled from Python, C++, Rust, Go, Zig, or another supported toolchain, with first-class ONNX inference.
Evaluate
The policy runs live across varied scenarios and opponents under a fixed wall-clock budget, every tick.
Understand
Scores, failure modes, per-tick compute, and synchronized 3D replay make the result inspectable.
Measured in the live run
342µsaverage agent time871µspeak agent time2,773ticks recorded
Sandboxed execution · ONNX inference · 3D physics · per-tick accounting · identical local and hosted engine · Gymnasium/PettingZoo packages · ladders · replays
Employer organizations · private campaigns · practice and evaluation suites · sealed submissions · comparative reports · retention and consent controls
Question 6 · Desired future state
Profitable in two years. The evaluation standard in five.
By year two, assessments, competitions, and turnkey environments support a six-person company without agent-evaluation revenue. By year five, Lockstep expands from hiring into continuous evaluation of production systems.
Assessment subscriptions plus a disciplined number of sponsored competitions and turnkey projects support the business without requiring agent-evaluation revenue.
Lockstep evaluates candidates, compares policy releases, qualifies third-party systems, and tests behavior before deployment across multiple realtime domains.
Assessments are a venture-scale company on their own: 10,000 accounts at a $5,000 mature blended annual value produce $50M in modeled revenue—less than 3% of today's cited assessment-software market. Competitions, turnkey environments, and recurring agent evaluation take the complete portfolio to $101M in revenue and $31.4M in operating profit. These are ambitions with explicit assumptions, not reported traction or a probability-weighted forecast.
Every score links back to a role requirement. The hiring team can see which scenario produced it, understand what the system did, and watch the replay—rather than trusting an unexplained ranking.
Question 7 · Where is the demand?
The market signal exists. Now we prove customers will pay.
Employers already pay for technical assessment and commit senior-engineering time to role-specific evaluation. That is the market demand Lockstep enters. The live evaluation core is built; the Fellowship funds the employer workflow, customer discovery, and first campaigns needed to convert that demand into Lockstep revenue.
Verified Market Research estimates the global pre-employment testing software market will grow at 7.02% annually. This is an established and expanding budget category; Lockstep begins with realtime-AI roles that general assessment platforms underserve. Market source ↗
Access to validate the thesis
Warm paths to buyer and practitioner conversations
- BigID VP of Data Science & Analytics Yehoshua Enuka has offered product counsel grounded in AI/ML hiring needs
- AI GTM specialist Moshe Porat has offered company introductions when the workflow is ready
- Founder-led outreach will target companies with multiple live physical-AI roles
- Direct Farama community access supports practitioner feedback and candidate reach
Question 8 · What are the economics?
At scale, about 75¢ of every assessment dollar remains after delivery costs.
The $990 and $4,500 plans stay close to the assessment platforms employers already buy. The first campaigns require hands-on setup and support. Once that repeated work becomes self-serve, the model projects roughly 25¢ in direct delivery costs for every $1 of assessment revenue.
Founder support covers one existing environment and up to ten candidates. It does not include a custom world. The annual plans only work financially when reports, candidate management, and routine support become self-serve.
The model's stress test: Lockstep would lose money if every attempt required the hands-on support planned for the first campaign. The founding offer therefore limits that support, and every repeated task becomes something the product must automate.
The downloadable model funds the founder plus three immediate hires at credible cash salaries, alongside the costs of building, selling, and operating the company. Other operating costs include $48K for infrastructure and tools, $45K for sales, travel, and events, $40K for legal, insurance, and administration, and $24K for specialist contractors.
- Cash salaries
- $625,000
- Payroll taxes and benefits
- $93,750
- Other operating costs
- $157,000
- Revenue-linked delivery costs
- $37,980
- Total 12-month spending
- $913,730
- 12-month revenue goal
- $116,920
- Cash remaining after 12 months
- $203,190
- Total time funded in this model
- ~14.8 months
Example that reaches profitability
$1.83M revenue · $50.6K left after operating costs600 blended assessment accounts + three sponsored competitions + six turnkey environments. This supports the modeled six-person team without requiring agent-evaluation revenue.
Question 8b · How does revenue compound?
Adoption builds the standard. The standard unlocks new markets.
Within each realtime domain, hiring teams and candidates benefit from a familiar evaluation platform instead of a bespoke process for every employer. As adoption spreads across roles and industries, Lockstep becomes the trusted assessment standard. That credibility attracts sponsors; sponsored competitions expand the audience and environment catalog; private suites turn that infrastructure into higher-value recurring evaluation.
Hiring teams converge on a trusted standard
A familiar candidate experience, comparable evidence, and recognized environments reduce the need for every employer to invent its own assessment.
More roles · more industriesFamiliarity lowers adoption friction
Candidates know how to submit their systems, reviewers know how to interpret the evidence, and each new domain broadens the reusable catalog.
Lower friction · broader adoptionSponsors want the standard
A concentrated practitioner audience and credible public benchmark give sponsors a reason to fund competitions, environments, and distribution.
Sponsor-funded expansionPublic proof creates private demand
Competition worlds demonstrate what is possible. Employers and agent teams then fund private adaptations, scenario suites, and repeated evaluation.
Turnkey + recurring evaluationThe highest-value expansion
A hiring suite becomes recurring evaluation infrastructure.
Once a team trusts a private suite to compare candidates, the same evaluation method can compare policy releases, qualify third-party systems, and expose regressions before deployment. That creates subscription-and-usage revenue each time the system changes. Agent evaluation is not required for profitability; it enters the model only after a team pays to rerun a private suite.
Catalog rule: confidential data and scenarios remain private; reusable primitives stay with Lockstep. Exclusivity is priced as the value of removing an asset from the catalog.
Question 9 · What is the wedge—and how does the model resolve?
Start with one paid assessment. Scale through reuse.
The first sale is a paid assessment for one active role using an existing environment: no custom world, no enterprise integration, and no claim that Lockstep replaces a general coding screen. First we prove the evidence matters. Then we turn repeated work into software.
Launch target: the first complete candidate assessment
For one active controls or policy role, the employer gets an approved rubric, candidate practice access, a private test configuration, replay-linked evidence, and a report review. Founder support covers up to ten candidates; the remaining annual attempts use the self-serve workflow.
One founder-supported assessment
Founder support connects the live runtime to the standards the employer uses for that role.
Repeated work becomes software
Rubrics, private test suites, reports, and candidate management become the assessment product—not ongoing consulting.
The next assessment launches faster
Each repeat role benefits from the employer's prior setup, reviewer familiarity, and growing evidence history.
Each new world opens another market
Competitions and paid environment work add reusable worlds; agent evaluation adds a higher-value recurring buyer.
AwsmRenderer ↗, Lockstep's open-source WebGPU renderer, supports MuJoCo and custom models. Once customer demand justifies a new environment, Lockstep can render an existing robotics model or build a relevant custom world without rebuilding the visualization stack.
One paid assessment proves the workflow. Reusable environments turn it into a platform.
Question 10 · Is this the business you were meant to build?
Lockstep is the company David spent twenty years preparing to build.
Lockstep combines sandboxed WebAssembly, realtime simulation, inference, metering, rendering, and adversarial system design. David Komer has shipped across that exact intersection—and will personally own the early customer relationships.
Founder & CEO
Twenty-plus years in engineering and more than eight as CTO across games, simulations, realtime systems, WebAssembly, and financial infrastructure.
Built JiCloud, one of the largest early WebAssembly platforms for browsers, in partnership with the creators of SQLx; then led work on WAVS under the inventor of CosmWasm.
Built AwsmRenderer ↗, a from-scratch open-source WebGPU renderer, before WebGPU was broadly available.
Built always-on crypto and financial infrastructure, including smart contracts that secured millions in total value locked.
Along with early work in WASM and WebGPU, built video streaming systems before YouTube and VR/AR systems before Oculus.
Competitions → assessment
The live public platform remains the distribution layer; recurring employer spend becomes the business wedge.
Gaming → real-world AI
The live arenas proved the engine. The first commercial wedge applies it to hiring people who build systems whose failures carry real-world consequences.
Composite score → inspectable evidence
Metrics and replays support a human decision instead of pretending to automate the hiring decision itself.
The first three hires
The platform is built. The round adds evaluation, environment design, and go-to-market.
BigID VP of Data Science & Analytics Yehoshua Enuka has tentatively agreed to advise implementation for equity.
- RL researcherBaselines, scenario design, and defensible evaluation.
- Simulation / environment designerCandidate-ready worlds and a growing catalog.
- GTM leadSupports a sales motion David continues to own.
Question 11 · Why build this with Gateway X?
The core technology is live. Build the company together.
Lockstep has built the execution engine beneath the product. The Fellowship combines initial capital with an embedded operating partnership in St. Louis. That is what Lockstep needs to discover the right assessment product with real hiring teams—and then build it quickly.
The execution engine
Sandboxed WASM and ONNX inference, realtime simulation, per-tick metering, scoring, ladders, and replay are live.
The assessment product
Private campaigns, candidate invitations, employer-approved rubrics, comparison reports, reviewer workflow, and data controls.
Operate like co-founders
Working side by side turns customer conversations, product decisions, and go-to-market iteration into shared operating work—not occasional advice.

Join the first buyer conversations, sharpen the offer from real objections, and turn founder selling into a disciplined weekly practice.

Turn what early customers need into the smallest useful assessment workflow—and separate repeatable software from custom work we should refuse.

Measure the funnel, launch time, support minutes, campaign margin, and hiring-team usage before adding complexity or headcount.
First buyer → evidence they trust → smallest paid offer → repeatable product
What we aim to accomplish together
- Identify the first buyer and role through direct customer discovery
- Define the evidence hiring teams will trust in a real decision
- Build the smallest complete employer assessment workflow
- Run paid campaigns and measure actual delivery economics
- Use the evidence to choose the environments, offer, and channel worth scaling
Extra · Competition · Where Lockstep fits
Lockstep does not need to replace the incumbents. It needs to own realtime-system evaluation.
These are the closest businesses competing for the budgets Lockstep eventually touches: technical hiring, simulation, sponsored competitions, and AI evaluation. The comparison is intentionally specific about where each company is stronger today—and the narrow market Lockstep can credibly win first.
| Comparable | What they do | Rough size ($) | My version at that scale | How we win market share |
|---|---|---|---|---|
| HackerRankAssessment subscriptions | Technical skill assessments and live interview software sold to employers through monthly and annual subscriptions. | Revenue not disclosed~$500M valuation in 2022; 2,500+ current business customers. No vetted public revenue figure found. | At 2,500 employer accounts, Lockstep's current blended plan model produces about $5.1M in annual assessment revenue. That makes assessment the entry point—not the entire outcome. | Do not fight its enormous generic question library. Own controls, autonomy, and policy roles where trajectories, recovery, latency, and resource use—not pass/fail code—determine quality. |
| CodeSignalSkills-based hiring | Assessments and interviews for technical and business hiring, with standardized tests, AI interviewers, proctoring, and enterprise workflow. | Revenue not disclosed$87.5M raised through its 2021 Series C. Current public plans are $948 and $5,748 per year, plus custom enterprise pricing. | At CodeSignal's breadth, Lockstep is the specialist assessment standard for systems that act, with the same recurring employer workflow extended into production-agent evaluation. | Win roles CodeSignal's static and conversational formats underserve. Show stateful behavior inside live worlds, then integrate with the employer's existing general assessment rather than demand a full replacement. |
| Applied IntuitionPhysical-AI simulation | Simulation, validation, and vehicle software for autonomous-system teams in automotive, defense, construction, and other physical industries. | ≈$400M ARRReported for year-end 2024 by The Information; private-company figure, not audited public revenue. | At this scale, private simulation and continuous production evaluation dominate. Hiring assessments remain distribution into a cross-industry physical-AI evaluation platform. | Applied Intuition owns deep enterprise vehicle workflows. Lockstep enters earlier and more horizontally through portable, untrusted-policy evaluation—starting with hiring, not a full autonomy toolchain. |
| KaggleSponsored AI competitions | Hosts public and private data-science and AI competitions for organizations, alongside a large practitioner community and free self-serve contests. | Revenue not disclosedGoogle-owned; 32M+ members and 33,000 competitions. Official materials show featured hosting fees of roughly $50K–$100K plus prizes. | At Kaggle's reach, sponsored seasons become a global acquisition channel and fund a large environment catalog that also serves hiring and agent evaluation. | Kaggle is strongest on datasets, prediction, and notebooks. Lockstep wins interactive, stateful competitions where policies must act continuously inside physics, adversaries, and realtime limits. |
| ArenaPublic-to-private AI evaluation | Runs a free public AI leaderboard and sells private evaluations and performance analysis to model labs and enterprises. | $100M annualized run rateReported in June 2026 after eight months of commercial sales; a consumption run rate, not full-year revenue or contracted ARR. | Public arenas create participation, data, and trust; private, usage-priced agent evaluation becomes the $100M-class engine. | Arena measures human preference over model outputs. Lockstep owns objective behavior under physics, timing, and compute constraints for agents that must control a changing system. |
Figures are the latest credible public disclosures located as of 18 August 2026. ARR, annualized run rate, full-year revenue, and valuation are deliberately not treated as interchangeable.
Reusing each environment turns a lower-priced hiring product into a platform. HackerRank and CodeSignal prove the buyer and budget. Applied Intuition proves physical-AI simulation can support a large enterprise business. Kaggle proves sponsors pay for technical competitions and distribution. Arena proves public evaluation can lead to paid private evaluation. Lockstep's opening is the shared execution layer none of them centers: running untrusted systems inside reusable realtime worlds.
Extra · Pseudo Jesse · Internal red team
The hard questions point directly to the twelve-week plan.
This is a deliberately hostile reading of the package—not a claim about Jesse's actual views. The criticisms below are not reasons to wait; they are the customer, product, and economic questions Lockstep and Gateway X can work through side by side.
Lockstep does not need twelve weeks to discover whether the core engine can be built; it needs them to turn that live engine into a buyer-shaped product and measurable company. Working side by side, we can talk directly with customers, ship the employer workflow, test paid campaigns, measure delivery economics, and learn what earns repeat use. Those outcomes directly answer the objections below.
What's weakThe definition is clear, but the live technology is not yet a candidate-ready employer product.
Why he'll pushA working engine can hide the product work that remains: invitations, rubrics, reports, comparison, privacy, and reviewer workflow.
What makes it solidOne end-to-end campaign demo: invite two candidates, run submissions, compare evidence, and deliver the report a hiring manager would actually use.
What's weakThe buyer and trigger are specific, but they still come from a thesis. Adviser access and promised introductions are not buyer evidence.
Why he'll pushA precise customer profile can still be imaginary if nobody confirms the current process, pain, budget owner, and cost of a bad screen.
What makes it solidFive recorded buyer interviews that agree on the role, current workaround, failure mode, purchase owner, and willingness to trial.
What's weakWASM, realtime execution, metering, and replay are distinct capabilities. The package has not proved employers value them together.
Why he'll pushA technical gap is not automatically a market gap. Buyers may accept lower fidelity because it is faster, familiar, and integrated.
What makes it solidA side-by-side role study where a conventional screen passes a submission that visibly fails Lockstep's live scenario—and the hiring manager says that difference changes the decision.
What's weakWebAssembly, ONNX, simulation, and AI-assisted coding did not all become true at once. The section risks assembling several trends instead of identifying one crossed threshold.
Why he'll pushIf the decisive change was available years ago, he will ask why this company wins now and why incumbents have not already shipped it.
What makes it solidOne dated before-and-after: the exact capability or cost threshold that made untrusted realtime evaluation practical, paired with current evidence that static screening is losing signal.
What's weakDance Off proves policies can run, compete, and replay. It does not yet prove assessment validity or employer usability.
Why he'll pushA compelling technical demo can still be far from a workflow someone budgets for and trusts in a hiring decision.
What makes it solidShow the same runtime producing an employer-approved scorecard, failure comparison, and replay-linked recommendation for a real role.
What's weakTwo-year profitability and five-year category leadership are goals from the model. The jump from hiring assessments to production-agent evaluation remains unproven.
Why he'll pushSimilar technology does not guarantee the same buyers, budgets, or validation standards in an adjacent market.
What makes it solidOne firm expansion rule: agent evaluation enters the plan only after an assessment customer pays to rerun a private test suite against production releases.
What's weakThere are no paying customers, live campaigns, signed pilots, or direct buyer quotes. Market activity is not Lockstep demand.
Why he'll pushThe company has already built substantial technology before proving the buyer will change process or pay.
What makes it solidThis cannot be fixed with copy. Keep naming it as the round's central risk; only a paid campaign and repeat use resolve it.
What's weakThe projected 74–76¢ remaining from each dollar depends on support time, environment calibration, report review, and candidate operations that have not yet been measured.
Why he'll pushA $990 plan can lose money if every employer needs custom scenarios or significant founder interpretation.
What makes it solidRun one campaign and publish the actual compute cost, setup hours, support minutes, report time, and amount left after direct delivery costs.
What's weakIndustry convergence, sponsor demand, and private evaluation form a credible sequence, but Lockstep has not yet observed a single handoff between them.
Why he'll pushA compounding diagram is not compounding revenue until adoption lowers the cost or raises the probability of the next sale.
What makes it solidMultiple employers adopt a familiar assessment without bespoke rebuilds, one sponsor funds a public environment, and one team pays to rerun a private suite across releases.
What's weakThe offer lowers purchase friction but may not cover founder-led sales, calibration, and support. The path from first campaign to meaningful account value is not observed.
Why he'll pushVolume is not a strategy until acquisition and service costs support it.
What makes it solidOne $990 campaign that converts into the $4,500 annual plan with materially less setup and support on the second role.
What's weakThe technical founder fit is unusually strong; the remaining gap is commercial capacity around a solo founder.
Why he'll pushBuilding the runtime and personally selling the first ten accounts are different jobs, both currently concentrated in David.
What makes it solidConvert tentative help into explicit commitments: signed adviser scope, named introduction targets, and a ready first-hire pipeline.
What's weakThe twelve-week plan names the work, but Gateway X still has to believe David will absorb uncomfortable customer feedback and change the product quickly.
Why he'll pushWorking like co-founders only matters if customer evidence can override the founder's assumptions.
What makes it solidA weekly scorecard for conversations, offer changes, product decisions, paid campaigns, launch time, support cost, and conversion—with explicit decisions to stop, change, or double down.
The three objections most likely to kill the meeting
- “You built the difficult technology before proving anyone will buy the workflow.”The honest answer is that this is the risk the round funds—not traction the company already has.
- “The $990 wedge cannot support founder-led sales and custom simulation work.”The company must show that repeat setup gets dramatically cheaper and that customers move into higher-value plans.
- “Realtime-AI assessment is a feature, not a venture-scale market.”HackerRank and CodeSignal prove assessment can support venture-backed platforms. Lockstep must prove its specialist wedge spans enough realtime industries to become a category—and that competitions, turnkey environments, and agent evaluation compound the core rather than compensate for it.
Measure behavior.
Build trust in realtime systems.