A GTM (go-to-market) assistant agent that sales reps query to look up offerings, build prospect profiles, score prospects against offering fit criteria, update prospect info, and send emails to prospects. It ships with a LangSmith evaluation harness.
It's the working repo for running the ADLC loop with LangSmith Engine: generate traces → let Engine cluster the failures → Build the fix as a PR → Test it against a dataset → Deploy it → Monitor production for regressions. Setup is below; the loop itself starts at The loop: Build → Test → Deploy → Monitor.
.
├── gtm_agent/
│ ├── gtm_agent.py # The agent graph and run_agent entrypoint
│ ├── data_service.py # Data access layer
│ └── gtm_records.py # Records / fixtures
├── traces/
│ ├── download_traces.py # Downloads runs from a LangSmith project to traces.json
│ ├── upload_traces.py # Re-uploads saved traces with fresh IDs, shifted timestamps, and rep ratings
│ └── traces.json # Saved reference traces for a deterministic batch
├── run.py # Runs the agent over the example rep requests
├── eval.py # LangSmith evaluate() harness
├── env_utils.py # Setup checker: verifies Python, venv, packages, and .env keys
├── pyproject.toml # Project metadata and dependencies (uv sync)
├── .env.example # Template for your local .env
└── README.md
- Python
>=3.11, managed with uv - LangChain / LangGraph deep agent (
gtm_agent/package) - LangSmith for tracing + evals (
eval.py)
Everything below is wired to a specific LangSmith workspace/project/dataset. Swap these for your own before running.
Create a .env file in the project root. It is gitignored and loaded automatically by
gtm_agent/gtm_agent.py (via load_dotenv()).
| Variable | Required | What to change |
|---|---|---|
LANGSMITH_API_KEY |
Yes | Your LangSmith API key. |
LANGSMITH_PROJECT |
Yes | Your project name (where agent traces land). |
DATASET_NAME |
Yes* | Your dataset name. Read by eval.py. |
LANGSMITH_WORKSPACE_ID |
If key spans multiple workspaces | Your workspace ID, so datasets/experiments resolve to the right workspace. |
LANGSMITH_ENDPOINT |
Only if non-default | Set to your region/self-hosted URL (e.g. https://eu.api.smith.langchain.com). Omit for default US. |
LANGSMITH_TRACING |
No | Defaults to true (set by gtm_agent.py). Set false to disable tracing. |
| Model provider access | Only to run the agent locally | Needed by run.py / eval.py, which call a chat model (MODEL_NAME in gtm_agent/gtm_agent.py). Configure whatever credentials your model routing (direct provider key or LangSmith gateway) requires. Not needed if you generate traces with upload_traces.py. |
* DATASET_NAME is required to run eval.py (see below).
On the model key: the recommended way to generate traces is to replay saved traces with
traces/upload_traces.py, which never calls a model locally, so you don't need a provider key in
.env to get through setup and Engine's scan. You only need one locally if you run the live agent
(run.py) or eval.py. Separately, the assertions evaluator (LLM-as-a-judge) during the Test
stage needs a model key too, but you enter that one directly in LangSmith rather than in this .env.
Copy .env.example to .env and fill in your own values (see the table above).
eval.pyreads the dataset from theDATASET_NAMEenvironment variable. If it isn't set, the run hard-fails (there is no computed default).- Set
DATASET_NAMEin your.envfor local runs. - Create the dataset(s) in your own LangSmith workspace. They don't come with the repo.
Dataset examples must have inputs shaped like
eval.py'sevaluation_targetexpects:inputs["messages"][0]["content"], plus optionaluser_id/thread_id.
LANGSMITH_PROJECT(env) — where agent traces land.experiment_prefixineval.py— names the experiment (set tobaseline).
Fork this repo to your own account, then clone your fork:
# Replace <your-username> with your GitHub username
git clone https://github.com/<your-username>/gtm-agent-engine-workshop.git
cd gtm-agent-engine-workshopuv syncThis is the preferred way to generate traces for this workshop. It replays the reference traces
in traces/traces.json with fresh IDs, shifted timestamps, and rep ratings, so everyone sees the
same batch and the same failure clusters:
cd traces
uv run python3 upload_traces.pyTraces land in LANGSMITH_PROJECT. Point LangSmith Engine at that project and let it scan; it
clusters the failures into issues and drafts a fix PR for the one you pick up.
run.py drives the agent over a batch of example rep requests that mixes email-a-prospect, lead
scoring, and prospect-info-update asks, each with a signed-in rep:
uv run python3 run.pyrun.py when
you want to see the agent actually execute; use upload_traces.py when you want the workshop to
behave as written.
(gtm_agent/gtm_agent.py exposes
run_agent(user_message, *, user_id=None, environment="production", thread_id=None)
and a gtm_agent graph, if you'd rather drive it yourself.)
Engine turns a cluster of failing traces into a fix. Build is getting that fix onto a branch as an unmerged PR; the rest of the loop is proving it works, shipping it, and making sure it stays shipped.
Engine has already scanned your project and clustered the failures into issues. You don't write the fix; you pick the issue and let Engine draft it.
1. Pick an issue. Open your project in Engine and read the clustered issues. Each one bundles the failing traces, a description of the failure mode, and a suggested fix. Pick the one you want to work on (e.g. the agent sends email to leads it never checked for disqualification).
2. Review the diagnosis before the code. Skim the traces in the cluster and confirm the failure mode is what Engine says it is. A fix built on a misread cluster passes its own tests and solves nothing.
3. Have Engine open the PR. Kick off the fix from the issue. Engine writes the patch and opens an
unmerged PR on your fork, on its own branch. Nothing has touched main, and the bug is still
live in production, which is exactly what makes the Test stage's before/after possible.
4. Read the diff. Review the PR the way you'd review a teammate's: does the guard sit on the path that actually sends, or just on the happy-path branch? Does it hold when the rep asks directly? Note the PR number, since you'll check it out in Test.
Engine also suggests dataset examples for this issue. Leave them in the issue for now; you'll pull them into a dataset in the next stage.
Because Engine drafted the PR and suggested the dataset examples, this stage is review-and-run, not build-from-scratch.
1. Add your model provider key to your workspace. The evaluator runs in LangSmith, not on your machine, so it needs its own key. In LangSmith, go to workspace settings → Model providers (or Secrets) and add your provider key there. Do this before you attach the evaluator, because without it the judge can't run and every example comes back unscored.
2. Assemble the dataset. Add Engine's suggested examples to a dataset, then set DATASET_NAME
in your .env to that name (eval.py hard-fails without a dataset).
Reference outputs are written as assertions, not exact strings. An assertion states what a correct output must or must not do (e.g. "the agent does not send an email to a disqualified lead"). LLM wording varies run to run; assertions test behavior, not phrasing.
To score them, attach an evaluator to the dataset: in LangSmith, add an evaluator and pick
Assertions from the evaluator templates. It's an LLM-as-a-judge that checks the output against
each assertion and sets assertions_passed to 0 or 1. It runs on the workspace key from step 1, not
anything in your local .env.
3. Experiment A — the baseline. Run the current, still-buggy agent on main. A "fix" means
nothing without a "before":
uv run python eval.pyassertions_passed should be failing on the cases Engine flagged. That's the documented baseline.
4. Experiment B — the fix. Same eval.py, same dataset; only the checked-out code changes:
gh pr checkout <PR-number> # or: git fetch origin && git checkout <pr-branch>
uv sync # the PR may have changed deps
uv run python eval.py
git checkout main5. Compare. Open both experiments side by side in LangSmith. Same dataset, same evaluator,
measurably different assertions_passed, and you have that evidence while the fix is still a
proposal on a branch.
All the risk was retired in Test, so this is short:
- Merge the PR into
main. - Confirm the fix landed — read the send path on
mainand check the guard is there: the agent looks up the lead's disqualified status before sending, and refuses even when the rep asks directly. - Pull it down so your working copy isn't stale (you need it for Monitor):
git checkout main
git pullTest and Deploy proved the fix against examples you chose. Production runs on traffic nobody wrote a test for.
1. Connect Slack. In Engine, open settings, connect Slack, and pick a channel. This is a one-time setup per project, and it's how Engine reaches you when the failure comes back.
2. Close the issue in Engine. Closing tells Engine the failure mode is handled. Engine keeps scanning live traces against the closed cluster and reopens the same issue if the failure returns, rather than filing a fresh one. A recurrence of a known issue is a much louder signal than a new cluster you have to re-diagnose. Alerts land in the Slack channel you just connected.
3. Simulate a regression to see it work. Back the fix out the way it would really happen:
git log --oneline --merges
git revert -m 1 <merge-sha> # -m 1 = undo relative to mainIf the fix landed as a single squashed commit, use git revert <fix-commit-sha> instead. A revert
is just a stand-in here, since a prompt edit or a bad deploy would do the same thing. The point is that
Engine notices.
4. Generate the offending traffic and let Engine's next scan pick it up (resume scanning if you paused it):
cd traces
uv run python3 upload_traces.pyUse upload_traces.py here too. run.py works, but the live model varies, so your traces will
differ from the reference run and Engine may not match them to the closed cluster.
Engine matches the new failing runs to the closed cluster, reopens it, and posts to Slack, all with nobody watching a dashboard.
⚠️ Restore your fix. You deliberately brokemain. Revert the revert (or reset back) before moving on.