
Marco-o1
🤖 Marco-o1: Towards Open Reasoning Models for Open-Ended SolutionsThe Alibaba MarcoPolo team presents Marco-o1, a large reasoning model designed to excel in open-ended problem-solving. Building upon ...
23 Marras 202414min

Test-Time Training
⌛️ The Surprising Effectiveness of Test-Time Training for Abstract ReasoningThis paper examines how test-time training (TTT) can enhance the abstract reasoning abilities of large language models (LLMs...
14 Marras 202414min

Qwen2.5-Coder
🔷 Qwen2.5-Coder Technical ReportThe report introduces the Qwen2.5-Coder series, which includes the Qwen2.5-Coder-1.5B and Qwen2.5-Coder-7B models. These models are specifically designed for coding ta...
12 Marras 202424min

Attacking Vision-Language Computer Agents via Pop-ups
😈 Attacking Vision-Language Computer Agents via Pop-upsThis research paper examines vulnerabilities in vision-language models (VLMs) that power autonomous agents performing computer tasks. The author...
9 Marras 202421min

Number Cookbook
📓 Number Cookbook: Number Understanding of Language Models and How to Improve ItThis research paper examines the numerical understanding and processing abilities (NUPA) of large language models (LLMs...
8 Marras 202416min

Jigsaw Puzzles
🧩 Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language ModelsThis research paper investigates the vulnerabilities of large language models (LLMs) to "jailbreak" attacks, where mali...
7 Marras 202416min

Multi-expert Prompting with LLMs
🤝 Multi-expert Prompting with LLMsThe research paper presents Multi-expert Prompting, a novel method for improving the reliability, safety, and usefulness of Large Language Models (LLMs). Multi-exper...
5 Marras 202412min



















