Eugene Siow

eugenesiow

AI & ML interests

None yet

Recent Activity

liked a model 2 days ago
google/gemma-7b-aps-it
liked a model 4 days ago
bytedance-research/UI-TARS-72B-DPO
liked a model 19 days ago
nomic-ai/modernbert-embed-base
View all activity

Organizations

Spaces-explorers's profile picture DSO National Laboratories's profile picture AI Singapore's profile picture DSO Intern's profile picture Top Contributors: Model Downloads's profile picture open/ acc's profile picture

eugenesiow's activity

liked a Space about 1 month ago
upvoted an article about 2 months ago
view article
Article

šŸŗšŸ¦ā€ā¬› LLM Comparison/Test: 25 SOTA LLMs (including QwQ) through 59 MMLU-Pro CS benchmark runs

By wolfram ā€¢
ā€¢ 76
reacted to m-ric's post with šŸ”„ about 2 months ago
view post
Post
1492
š—¦š—µš—¼š˜„š—Øš—œ: š—® š˜€š—ŗš—®š—¹š—¹ š—²š—»š—±-š˜š—¼-š—²š—»š—± š—®š—“š—²š—»š˜ š˜š—µš—®š˜ š—°š—®š—» š—»š—®š˜ƒš—¶š—“š—®š˜š—² š—®š—»š˜† š—Øš—œ š—®š—»š—± š—¼š˜‚š˜š—½š—²š—暝—³š—¼š—暝—ŗš˜€ š—ŗš˜‚š—°š—µ š—Æš—¶š—“š—“š—²š—æ š˜€š˜†š˜€š˜š—²š—ŗš˜€! šŸ“²

A team from NUS and Microsoft just released an agent that can act on any UI (Desktop, Android, Web) without needing additional text information. It works extremely well : they applied their method on a tiny Qwen2-VL-2B, and they managed to beat methods that use either much more powerful vision models (like GPT-4V) without using any additional info (e.g. leveraging the DOM of a webpage) like previous methods did ! šŸ‘šŸ‘

They started from the idea that most existing methods rely heavily on text, which makes them less generalizable, while letting aside rich UI structure that user actually rely on when navigating this interfaces.

āš™ļø They put several good ideas to work:

šŸ’” Simplify screenshots to the max:
They prune a lot the heavy visual content of UI screenshots, by removing cloned image patches (like any vast patch of the same color will be reduced to a small patch, while maintaining positional embeddings), then group patches from the same GUI elements together to simplify even further

šŸ’” Build a truly generalist dataset:
To train a general UI agent, you need trajectories from each possible UI, and express them in a common language. Authors merge datasets like OmniAct for Desktop, Mind2Web for websites, AMEX for Android trajectories to create a high-quality and diverse dataset.

āž”ļø Nice results ensued:
They fine-tune a tiny Qwen-2-VL-2B on their method, and it reaches SOTA on several task (element identification, web navigation), even beating methods that either use additional info from the DOM or use much bigger VLMS like GPT-4v! šŸ†

And performance could certainly jump with a slightly bigger vision model. Let's hope the community builds this soon! šŸš€

Paper added to my "Agents" collection šŸ‘‰ m-ric/agents-65ba776fbd9e29f771c07d4e