PolySearchRESEARCH

ONGOING RESEARCH · TARGETING SIGIR 2027

PolySearchBehavior-Aware
AI Tool Search

Learning Reusable Tool Behavior
from Real Interactions

PolySearch investigates how interactive search systems can recommend AI tools for evolving user needs. We explore execution-grounded behavioral representations that capture how tools perform across tasks and constraints.

当需求不断变化,工具推荐能否从真实交互中学习?我们研究如何复用历史行为证据,让推荐超越静态描述,减少重复试用。

Feasibility stage · No accepted paper or validated predictor claimed

THE RESEARCH DIRECTIONProposed
EVOLVING USER NEED

“Find evidence I can trace.”

Source-onlyPrecise citationsKnow when to abstain
↓ match task demands
REUSABLE BEHAVIORAL REPRESENTATION
CorrectnessGroundingRestraint
Task × toolObserved outcomeEvidence

Historical executions → task-conditioned evidence

↓ predict suitability
↗
Recommend with evidenceProbe selectively when history is insufficient

Conceptual architecture. Predictive benefit remains untested.

INITIAL DOMAINAcademic & Literature ResearchFrom Static Descriptions to Execution-Grounded Tool Representations

01 / RESEARCH MOTIVATION

Descriptions tell us what.
Interactions reveal how.

工具说明不等于任务表现。我们的目标是积累可迁移的行为证据,而不是每次从零测试所有工具。

01

Traditional retrieval

User query → Tool description → Ranking

Match a request against static documents.

LIMIT: ADVERTISED CAPABILITY ≠ OBSERVED PERFORMANCE
02

Execution-aware retrieval

Candidates → Live probing → Reranking

Gather direct evidence, at repeated execution cost.

LIMIT: ACCESS, LATENCY AND REPEATED PROBES

02 / RESEARCH QUESTIONS

Can tool history become reusable knowledge?

RQ1

Behavioral prediction

Can execution history predict performance on unseen academic tasks?

历史行为能否预测新任务中的表现?

RQ2

Behavior-aware retrieval

Can behavioral matching improve ranking beyond descriptions and capabilities?

行为证据能否改善工具检索与排序?

RQ3

Probe efficiency

Can we reduce fresh probing while maintaining recommendation quality?

能否减少重复试用,同时保持推荐质量?

ADescription onlyBStatic capabilitiesCAggregate historyDTask-conditioned history

The critical comparison: D > C on held-out source groups. These are planned comparisons, not measured results. Exact-probe caching alone is not generalization.

03 / TASK SCENARIOS

One domain. Different demands.

Academic research offers public benchmark sources, reference-grounded evaluation, observable citation behavior and multi-document workflows.

04 / BENCHMARK FOUNDATIONS

Reuse sources. Collect new behavioral evidence.

我们复用已有问题与文献来源,并制定任务级原子评分;新增的是实际工具执行记录,而不是宣称已经建立全新的基准。

Official paper/data links identify the source of each summary. No benchmark gold answers or private annotations are republished. Dataset licensing does not automatically grant rights to every publisher PDF.

05 / THIRD-PARTY TOOL LANDSCAPE

Real tools. Clearly scoped evidence.

Only ChatPDF Web and SciSpace Web were evaluated in this pilot. Other cards describe documented workflows, not demonstrated competence.

06 / INITIAL FEASIBILITY

What we actually observed.

Evidence snapshot

Loading verified data…

Loading review status…

Execution outcomes

All attempted executions · operational outcome only

Semantics and format are different axes

Observed outputs only. Unequal task coverage prevents an overall tool ranking.

A format difference, not a demonstrated semantic winner.

Both tools gave substantively correct answers on the paired CG tasks. ChatPDF met the strict JSON contract; SciSpace's responses failed the complete-output format contract. SciSpace also answered both PU tasks correctly.

ChatPDF 的 PU 任务因每日额度限制而没有答案,不能记为语义错误。当前证据无法证明任务相关的语义优劣;表格与 JSON 要求又与任务族部分混杂。

Four tasks, one initial attempt per tool/task. No overall winner; no validated historical predictor.

07 / INTERACTIVE EVIDENCE EXPLORER

Every cell has a record.

Choose a tool × task cell to inspect operational status, atomic outcomes and provenance. Missing evidence stays missing.

✓ All essential criteria passed△ Output contract failed⊘ Quota/access blocked○ Not executed— Not evaluated

Authentic evidence, bounded disclosure

Raw responses and screenshots exist in the private evidence store. They are not republished here because third-party content/UI redistribution permission has not been established.

公开图表来自真实记录;原始截图与文本未获确认的公开使用许可,因此保留为不可用状态,不以合成图片替代。

Public provenance & image exclusions ↗

Select an observation

Each cell links to its recorded evaluation.

08 / EXPERIMENT TIMELINE

Replication before representation.

先排除额度与顺序混杂,再判断差异是否稳定存在。计划日期不代表已经执行。

Prepared · October 6

Sources and tasks frozen

QASPER / SciFact source packets, task manifests and atomic criteria prepared.

Recorded · October 10

Calibration and pilot

Independent AI calibration reviewed 24 controls. Eight initial attempts recorded; human audit remains pending.

Prepared · October 10

Balanced replication

Fresh conversations, no answer repair, two quota positions and reversed task order across repetitions.

Balanced replication progress

Loading…

Each collection day: two tasks per tool, 4 p.m. New York time, subject to quota/access readiness. No out-of-window attempts without approval. Static export—not a live third-party connection.

09 / NEXT STEPS

A staged research agenda.

01
Completed pilot

Initial feasibility

Domain, benchmark exploration, frozen source packets and eight attempts.

02
Current · prepared

Balanced replication

Check stability, quota confounding, semantic ceilings and format effects.

03
Conditional

Source-disjoint expansion

New tasks only if meaningful behavior variation emerges. Near-ceiling semantics with format-only differences call for task redesign.

04
Proposed

Behavioral representation

Compare descriptions, static metadata, aggregate history and conditioned history on unseen tasks.

05
Proposed

Selective probing

Measure held-out quality against fresh execution cost, including the cost of building history.

10 / REPRODUCIBILITY & LIMITATIONS

Make the uncertainty inspectable.

What we preserve

  • Frozen sources and task manifests, with source hashes
  • Immutable observations and versioned atomic judgments
  • Independent, blinded AI review; human audit explicitly pending
  • Timestamped captures and recorded interventions
  • Separate operational and semantic outcomes

What this evidence cannot establish

  • Stable capability across externally changing, stochastic tools
  • Equivalent access across account tiers and quotas
  • Absence of hidden retrieval or benchmark contamination
  • Error-free rubric judgments or population-level generalization
  • Semantic superiority inferred from format adherence alone

Research decision after replication: EXPAND · REDESIGN TASKS · FIX COLLECTION. No automatic expansion, representation model or validated retrieval system is implied.