fix(typesafe): raise TsToolSelectorMiddleware default relevance_threshold to 0.5

0.3 kept nearly the entire tool roster on every task in a homogeneous
eval batch (contextbench lookup+write-answer tasks), giving the
filter little precision. Testing against a mixed-tool-requirement
batch showed 0.5 correctly drops clearly-irrelevant tools (e.g.
write_file/execute for a read-only "explain this function" query)
without cutting anything genuinely needed, while 0.6 already cuts a
borderline-relevant tool (grep at 0.59) in the same batch.
This commit is contained in:
thushanth-bengre-langchain committed 2026-09-18 14:07:27 -04:00
1 parent 4dd1449ec2
commit e9ac94d9f6
2 files changed
+2 -2

No files matched your search

+1 -1
View File
@@ -87,7 +87,7 @@ agent = create_agent(
)
```
The middleware asks one independent `Noul` question per candidate tool ("is this tool needed next?"), batched into a single TypeSafe request against the latest human message, before every model call. Tools whose probability clears `relevance_threshold` (default `0.3`) are kept, ranked by that probability, and capped at `max_tools` if set. Use `always_include` to keep specific tools regardless of classification. This API is experimental and may change without notice.
The middleware asks one independent `Noul` question per candidate tool ("is this tool needed next?"), batched into a single TypeSafe request against the latest human message, before every model call. Tools whose probability clears `relevance_threshold` (default `0.5`) are kept, ranked by that probability, and capped at `max_tools` if set. Use `always_include` to keep specific tools regardless of classification. This API is experimental and may change without notice.
### LangChain messages as state
@@ -90,7 +90,7 @@ class TsToolSelectorMiddleware(
def __init__(
self,
*,
relevance_threshold: float = 0.3,
relevance_threshold: float = 0.5,
max_tools: int | None = None,
always_include: list[str] | None = None,
) -> None: