Traditional retrieval
User query → Tool description → Ranking
Match a request against static documents.
LIMIT: ADVERTISED CAPABILITY ≠ OBSERVED PERFORMANCEONGOING RESEARCH · TARGETING SIGIR 2027
Learning Reusable Tool Behavior
from Real Interactions
PolySearch investigates how interactive search systems can recommend AI tools for evolving user needs. We explore execution-grounded behavioral representations that capture how tools perform across tasks and constraints.
当需求不断变化,工具推荐能否从真实交互中学习?我们研究如何复用历史行为证据,让推荐超越静态描述,减少重复试用。
Feasibility stage · No accepted paper or validated predictor claimed
“Find evidence I can trace.”
Historical executions → task-conditioned evidence
Conceptual architecture. Predictive benefit remains untested.
01 / RESEARCH MOTIVATION
工具说明不等于任务表现。我们的目标是积累可迁移的行为证据,而不是每次从零测试所有工具。
User query → Tool description → Ranking
Match a request against static documents.
LIMIT: ADVERTISED CAPABILITY ≠ OBSERVED PERFORMANCECandidates → Live probing → Reranking
Gather direct evidence, at repeated execution cost.
LIMIT: ACCESS, LATENCY AND REPEATED PROBESHistory → Representation → Behavioral matching
Reuse evidence across evolving needs. Selective probes add new observations when coverage is insufficient.
HYPOTHESIS: TRANSFER BEYOND THE ORIGINAL TASK02 / RESEARCH QUESTIONS
Can execution history predict performance on unseen academic tasks?
历史行为能否预测新任务中的表现?
Can behavioral matching improve ranking beyond descriptions and capabilities?
行为证据能否改善工具检索与排序?
Can we reduce fresh probing while maintaining recommendation quality?
能否减少重复试用,同时保持推荐质量?
The critical comparison: D > C on held-out source groups. These are planned comparisons, not measured results. Exact-probe caching alone is not generalization.
03 / TASK SCENARIOS
Academic research offers public benchmark sources, reference-grounded evaluation, observable citation behavior and multi-document workflows.
04 / BENCHMARK FOUNDATIONS
我们复用已有问题与文献来源,并制定任务级原子评分;新增的是实际工具执行记录,而不是宣称已经建立全新的基准。
Official paper/data links identify the source of each summary. No benchmark gold answers or private annotations are republished. Dataset licensing does not automatically grant rights to every publisher PDF.
05 / THIRD-PARTY TOOL LANDSCAPE
Only ChatPDF Web and SciSpace Web were evaluated in this pilot. Other cards describe documented workflows, not demonstrated competence.
06 / INITIAL FEASIBILITY
Loading verified data…
All attempted executions · operational outcome only
Observed outputs only. Unequal task coverage prevents an overall tool ranking.
Both tools gave substantively correct answers on the paired CG tasks. ChatPDF met the strict JSON contract; SciSpace's responses failed the complete-output format contract. SciSpace also answered both PU tasks correctly.
ChatPDF 的 PU 任务因每日额度限制而没有答案,不能记为语义错误。当前证据无法证明任务相关的语义优劣;表格与 JSON 要求又与任务族部分混杂。
Four tasks, one initial attempt per tool/task. No overall winner; no validated historical predictor.
07 / INTERACTIVE EVIDENCE EXPLORER
Choose a tool × task cell to inspect operational status, atomic outcomes and provenance. Missing evidence stays missing.
Raw responses and screenshots exist in the private evidence store. They are not republished here because third-party content/UI redistribution permission has not been established.
公开图表来自真实记录;原始截图与文本未获确认的公开使用许可,因此保留为不可用状态,不以合成图片替代。
Public provenance & image exclusions ↗Each cell links to its recorded evaluation.
08 / EXPERIMENT TIMELINE
先排除额度与顺序混杂,再判断差异是否稳定存在。计划日期不代表已经执行。
QASPER / SciFact source packets, task manifests and atomic criteria prepared.
Independent AI calibration reviewed 24 controls. Eight initial attempts recorded; human audit remains pending.
Fresh conversations, no answer repair, two quota positions and reversed task order across repetitions.
Each collection day: two tasks per tool, 4 p.m. New York time, subject to quota/access readiness. No out-of-window attempts without approval. Static export—not a live third-party connection.
09 / NEXT STEPS
Domain, benchmark exploration, frozen source packets and eight attempts.
Check stability, quota confounding, semantic ceilings and format effects.
New tasks only if meaningful behavior variation emerges. Near-ceiling semantics with format-only differences call for task redesign.
Compare descriptions, static metadata, aggregate history and conditioned history on unseen tasks.
Measure held-out quality against fresh execution cost, including the cost of building history.
10 / REPRODUCIBILITY & LIMITATIONS
Research decision after replication: EXPAND · REDESIGN TASKS · FIX COLLECTION. No automatic expansion, representation model or validated retrieval system is implied.