Abstract
Background. REST APIs have become the primary interaction interface for modern software systems, making automated black-box API testing increasingly important. However, existing testing techniques still face three major limitations: over-reliance on dependencies extracted from static specifications, failure to fully utilise failed requests as feedback, and an overemphasis on crash-oriented verification, resulting in insufficient exploration of non-crash-prone logical defects.
Aims. We design, implement and evaluate a feedback-driven black-box testing framework that turns runtime feedback from a passive outcome into an active control signal, maximising logical-defect yield within practical budgets.
Method. We propose FDRRestTest, which integrates OpenAPI-based static semantic bootstrapping, runtime dependency evolution, a 7-category failure diagnosis and repair loop, a multi-dimensional logical oracle engine and a budget-aware utility scheduler in one closed loop. Under a fixed 600-second time budget, we evaluated FDRRestTest against five state-of-the-art baselines (RESTler, EvoMaster, Morest, ARAT-RL, AutoRestTest) on 12 real-world REST services along five protocols: fixed-time cross-tool comparison, fixed-request cross-tool replication, ablation across three dependency profiles, single-host budget sensitivity, and per-service comparison; significance via two-sided paired Wilcoxon signed-rank tests with Holm-Bonferroni correction, magnitude via standardised paired effect sizes.
Results. FDRRestTest improves fault-revealing efficiency (FRE) by up to 6.3% over the strongest baseline, AutoRestTest (0.9878 vs. 0.9294; an upper estimate, as AutoRestTest’s FRE is a post-hoc lower bound), while issuing 6.3-11.8× fewer requests than AutoRestTest across the 12 services (8.3× on Features Service). The +12.2% logical-yield advantage holds on every one of the 12 services (+7.0%-+16.5%, geometric mean +12.4%), and the FRE gain is consistent across all 12 in per-service means and statistically significant (p_{adj} < .005, d_z = 3.73 on logical bug yield, LBY). An ablation across six subjects spanning three dependency profiles confirms no single component accounts for the advantage, the oracle contributing most - removing it drops FRE below the strongest baseline on all six - while a Pareto-scheduling variant is 4.4% below the weighted utility on the most state-rich host. A manual confirmation study (N = 120, stratified over the 236 deduplicated findings of all six tools, 30 per oracle category) confirms 96/120 (80%) as true defects (Cohen’s κ = 0.78), supporting the oracle-finding interpretation of the LBY gain.
Conclusions. Black-box REST API testing heavily benefits from a tightly coupled closed feedback loop. Within this loop, runtime dependency evolution, failure repair, logical-defect-aware validation, and budget-aware scheduling are each necessary; quantifying their interaction would require multi-component removals, which we leave to future work.