跳到正文

jinzijian

EvoTrace

Compile real-world Claude Code and Codex trajectories into verified, tradable post-training assets.

README 已保存到本站,可直接阅读

Documentation snapshot

README 快照

本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。

EvoTrace

Turn real-world Claude Code and Codex sessions into reusable training, evaluation, and verification assets.

A local-first trajectory compiler built on DeepSeek Harness.

Get started · What you get · Workflow · Architecture · 中文

图片:CI 图片:DeepSeek Harness 图片:License: Apache-2.0

EvoTrace imports the coding-agent work already stored on your machine, finds the sessions worth keeping, and compiles them into evidence-backed preference data, replayable coding tasks, RL environments, and execution-reward candidates. You keep the source data and the resulting assets under your control.

It is not another coding agent and does not require you to change how you use Claude Code or Codex.

[!WARNING] EvoTrace is early alpha, and DeepSeek Harness is a developer preview. Generated tasks and verifiers remain candidates until they pass the documented evidence and Docker validation gates.

Get started

1. Install

macOS, Linux, or WSL:

curl -LsSf https://raw.githubusercontent.com/jinzijian/EvoTrace/main/install.sh | sh

Windows PowerShell:

irm https://raw.githubusercontent.com/jinzijian/EvoTrace/main/install.ps1 | iex

2. Launch

evotrace

EvoTrace opens the DeepSeek Harness Web app. In Settings, choose any provider supported by your Harness setup, such as DeepSeek, OpenAI, or Anthropic. Local import and deterministic mining do not require a model; agent review, hardening, calibration, and evolution do.

New sessions start with Harness Full access by default. To use the narrower workspace sandbox:

DSH_PERMISSION_MODE=workspace-write evotrace

Docker is required only when you build or validate executable environments.

3. Build your first asset

Type / in the app and run:

/init              import existing Claude Code and Codex history
/candidates        show the strongest evidence-backed sessions
/show 1            inspect candidate 1 and its missing evidence
/review 1          run the sequential four-agent review
/assets            inspect anything that was compiled or verified

That is the main product loop. A review may build an asset, route the session to preference data, keep it only as a hardening seed, or reject it with explicit reasons. Rejection is a useful result: it prevents a long but weak trajectory from being mislabeled as training-ready.

What you get

OutputRecovered or generated from your sessionsUseful for
Candidate catalogtask intent, repo, corrections, failures, effective actions, provenance gapsfinding the small fraction of history worth keeping
Preference and recovery datarejected/chosen attempts, human corrections, successful recoveriesDPO, SFT, QA, failure-recovery training
Executable task bundlerepository base, initial state, dependency evidence, task specificationcoding-agent evals, regression tasks, RL environments
Verifier and reward candidatetest commands, behavioral checks, policy, provenancerollout scoring and execution rewards after validation
Difficulty evidencefresh independent solver attempts and verifier outcomescurriculum construction instead of guessing from patch size
Execution experiencegrounded runtime facts compressed from exploration trajectoriestraining examples and held-out experience-transfer experiments

The same validated task can evaluate an agent today, score newly sampled rollouts tomorrow, and produce verifier-grounded RL data later. The future opt-in EvoTrace Marketplace and fine-tuning integrations are intended to let users license reviewed assets on terms they control; they are roadmap products, not part of the current local release.

Why raw trajectories are not enough

A transcript may contain a prompt, messages, commands, and diffs, but post-training needs more:

  • one coherent task boundary rather than an entire chat;
  • the repository state from before the task began;
  • a reproducible dependency and execution environment;
  • an independent verifier that rejects the base state and accepts a known-good state;
  • provenance tying every task, patch, verifier, and run together;
  • difficulty measured by fresh attempts rather than token count or patch size.

EvoTrace automates that compilation gap with session import, Git/repository archaeology, deterministic gates, specialized agents, and isolated execution.

The core workflow

Claude Code / Codex history
           │
           ▼
        /init         import + normalize + Git archaeology
           │
           ▼
     /candidates      rank evidence, hide nested subagent duplicates
           │
           ▼
       /review        mine episode → gate route → build/harden → criticize
           │
      ┌────┴───────────────┐
      ▼                    ▼
preference/recovery   executable candidate
                           │
                           ▼
                    /validate in Docker
                           │
                           ▼
                 verified reward environment

Status means evidence

StatusWhat it actually means
MinedThe session has useful signals. Nothing executable is implied.
BuildableTask, repo base, reconstruction confidence, reference patch, verifier commands, and environment gates pass.
Bundle generatedA Docker-ready candidate exists. Its verifier is not yet trusted.
VerifiedA conforming Docker run rejected the base, accepted the reference, and was recorded against the exact bundle digest.
CalibratedFresh solver attempts measured the task; the default target is two verifier passes in five attempts.

EvoTrace fails closed on empty prompt wrappers, low-confidence reconstruction, missing reference patches, missing verification commands, unsupported environments, candidate switching, and mismatched asset lineage.

Common recipes

Mine useful history without sending it to a model

/init
/candidates
/show 1

Compile and independently validate an executable task

/review 1
/build 1
/validate 1
/runs

Make an easy verified task meaningfully harder

/harden 1
/calibrate 2

Hardening must add testable behavior, compatibility, edge cases, or failure constraints. Making a patch longer is not treated as making a task harder.

Test whether execution experience transfers

/evolve 1 2

Asset 1 is explored and compressed; asset 2 must be an independently built held-out task from the same repository. Baseline and conditioned solver attempts are then compared using Docker rewards. Running /evolve 1 without a held-out asset is only a wiring smoke test and cannot certify transfer.

Command reference

CommandPurpose
/init [all|codex|claude]import history and refresh mining
/candidatesbrowse ranked candidates
/search payment retrysearch tasks, repositories, and evidence
/show 1inspect provenance and readiness gaps
/review 1run the sequential review pipeline
/build 1compile an execution candidate
/validate 1run two-state Docker validation
/harden 1derive and test a harder child task
/calibrate 1measure and adapt difficulty with self-play
/evolve 1 2test compressed experience on a held-out task
/assetslist compiled assets and their states
/runsinspect saved validation evidence
/doctorcheck local integrations

Built on DeepSeek Harness

EvoTrace is a specialized distribution of DeepSeek Harness. Harness supplies the Web UI, sessions, streaming, provider settings, permission surface, slash commands, and plugin runtime. EvoTrace adds the trajectory compiler and one managed Orchestrator with four foreground, least-privilege roles:

StageResponsibilityCannot do
Episode Minerisolate one coherent episode and count effective actionsbuild or approve an asset
Candidate Gatejudge value, complexity, reconstructability, and record one immutable routemutate data or change the route later
Task Builder / Hardenerbuild the exact routed candidate or derive a harder childedit the source checkout or approve itself
Verifier Criticaudit Docker runs, verifier evidence, lineage, and difficultycertify missing evidence

The children run sequentially, never in parallel. Review-bound tools enforce the exact review token, candidate ID, route, and produced-asset lineage in code rather than relying only on prompts.

Install from source

Requires Git, Node.js 22.19+ or 24+, Python 3.9+, and optionally Docker.

git clone https://github.com/jinzijian/EvoTrace.git
cd EvoTrace
python3 -m venv .venv
.venv/bin/python -m pip install -e .
pnpm install
pnpm dev

The Python CLI remains available as a deterministic compiler and automation sidecar. Run et --help for its machine-oriented commands; the DeepSeek Harness app is the primary interface.

Execution and trust boundaries

  • Import and mining read local Claude Code/Codex history and Git evidence without changing source repositories.
  • Codex subagent and fork trajectories keep parent lineage but are hidden from the default candidate list.
  • The Orchestrator exposes fixed domain tools rather than arbitrary host shell or filesystem tools.
  • Validation runs in disposable Docker worlds without source bind mounts, the Docker socket, host networking, privileged mode, or host credentials.
  • Builder and Verifier Critic are separate child sessions; a builder cannot approve its own verifier.
  • Self-play and evolution are explicit operations because they send selected task context to the configured model.

Read the normative sandbox contract, task quality standard, schema, and design.

Current release and roadmap

The current open-source release includes local Claude Code/Codex import, evidence mining, sequential agent review, environment reconstruction, Docker bundle generation, two-state validation, self-play calibration, semantic task hardening, experience compression, and held-out transfer measurement.

Still in progress:

  • broader cross-language dependency repair and autonomous environment construction;
  • stronger hidden behavioral verifiers and adversarial task mutation;
  • validated DPO, SFT, and RL dataset exporters;
  • opt-in Marketplace and managed fine-tuning integrations.

Acknowledgements

  • DeepSeek Harness is the application and agent foundation.
  • Microsoft RepoLaunch is a primary inspiration for reproducible repository-to-environment construction. EvoTrace begins earlier by mining tasks and learning signals from lived coding-agent work.

License

Apache-2.0. See LICENSE and NOTICE.

Official distribution

获取与安装

暂未发现可确认的官方软件包地址

当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。

本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。

使用前核验

本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。