Search Results

Documents authored by Jiang, Nanxiang


Document
Technical Track Paper
Can LLM Coding Assistants Support Emerging Programming Languages? An Empirical Study on Cangjie

Authors: Junhang Cheng, Fang Liu, Jia Li, Chengru Wu, Nanxiang Jiang, and Li Zhang

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. LLM coding assistants are widely used in development, but evidence about their reliability mostly comes from mainstream languages such as Python and Java. Their behaviour on emerging languages, where public code and documentation are scarce, is far less clear. Aims. We ask how capable current LLMs are on Cangjie, an emerging language in the HarmonyOS ecosystem, and which assistive technique is worth its token cost. Method. We build CangjieBench, 248 manually translated tasks derived from HumanEval and ClassEval, run under Text-to-Code and Code-to-Code modalities. We evaluate six LLMs under five methods: direct generation, syntax-constrained prompting, retrieval over documentation, retrieval over code, and agent-based workflows. We report Pass@1, compile rate, token cost, and a seven-category failure taxonomy validated by two LLM annotators (κ ≥ 0.98). Results. Direct generation is unusable on most models, and over four-fifths of failed samples carry foreign syntax that the Cangjie compiler rejects. A 2,146-token syntax reference gives the best accuracy-to-cost trade-off among prompt-based methods. The best agent configuration reaches the highest Pass@1 at 10-120x the token cost of a single prompt, while weaker agents cannot match a syntax-constrained call on the same backbone. A Python reference does not always help: on Code-to-Code it can pull models toward Python idioms and drop compile rates below Text-to-Code. Conclusions. No single method dominates both accuracy and token cost. A practical assistant for an emerging language should ship a concise syntax cheat sheet first, escalate to curated retrieval for API-heavy tasks, and reserve agent loops for tasks that need multi-turn repair. The replication package is at https://github.com/cjhCoder7/CangjieBench.

Cite as

Junhang Cheng, Fang Liu, Jia Li, Chengru Wu, Nanxiang Jiang, and Li Zhang. Can LLM Coding Assistants Support Emerging Programming Languages? An Empirical Study on Cangjie. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 44:1-44:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{cheng_et_al:LIPIcs.ESEM.2026.44,
  author =	{Cheng, Junhang and Liu, Fang and Li, Jia and Wu, Chengru and Jiang, Nanxiang and Zhang, Li},
  title =	{{Can LLM Coding Assistants Support Emerging Programming Languages? An Empirical Study on Cangjie}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{44:1--44:21},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.44},
  URN =		{urn:nbn:de:0030-drops-280128},
  doi =		{10.4230/LIPIcs.ESEM.2026.44},
  annote =	{Keywords: Large language models, Code generation, Code translation, Emerging programming languages, Empirical software engineering, Benchmark, Cangjie}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail