Tuesday, 13 May 2025
  • My Feed
  • My Interests
  • My Saves
  • History
  • Blog
Subscribe
Capernaum
  • Finance
    • Cryptocurrency
    • Stock Market
    • Real Estate
  • Lifestyle
    • Travel
    • Fashion
    • Cook
  • Technology
    • AI
    • Data Science
    • Machine Learning
  • Health
    HealthShow More
    Skincare as You Age Infographic
    Skincare as You Age Infographic

    When I dove into the scientific research for my book How Not…

    By capernaum
    Treating Fatty Liver Disease with Diet 
    Treating Fatty Liver Disease with Diet 

    What are the three sources of liver fat in fatty liver disease,…

    By capernaum
    Bird Flu: Emergence, Dangers, and Preventive Measures

    In the United States in January 2025 alone, approximately 20 million commercially-raised…

    By capernaum
    Inhospitable Hospital Food 
    Inhospitable Hospital Food 

    What do hospitals have to say for themselves about serving meals that…

    By capernaum
    Gaming the System: Cardiologists, Heart Stents, and Upcoding 
    Gaming the System: Cardiologists, Heart Stents, and Upcoding 

    Cardiologists can criminally game the system by telling patients they have much…

    By capernaum
  • Sport
  • 🔥
  • Cryptocurrency
  • Data Science
  • Travel
  • Real Estate
  • AI
  • Technology
  • Machine Learning
  • Stock Market
  • Finance
  • Fashion
Font ResizerAa
CapernaumCapernaum
  • My Saves
  • My Interests
  • My Feed
  • History
  • Travel
  • Health
  • Technology
Search
  • Pages
    • Home
    • Blog Index
    • Contact Us
    • Search Page
    • 404 Page
  • Personalized
    • My Feed
    • My Saves
    • My Interests
    • History
  • Categories
    • Technology
    • Travel
    • Health
Have an existing account? Sign In
Follow US
© 2022 Foxiz News Network. Ruby Design Company. All Rights Reserved.
Home » Blog » TokenSet: A Dynamic Set-Based Framework for Semantic-Aware Visual Representation
AITechnology

TokenSet: A Dynamic Set-Based Framework for Semantic-Aware Visual Representation

capernaum
Last updated: 2025-03-25 04:53
capernaum
Share
TokenSet: A Dynamic Set-Based Framework for Semantic-Aware Visual Representation
SHARE

Visual generation frameworks follow a two-stage approach: first compressing visual signals into latent representations and then modeling the low-dimensional distributions. However, conventional tokenization methods apply uniform spatial compression ratios regardless of the semantic complexity of different regions within an image. For instance, in a beach photo, the simple sky region receives the same representational capacity as the semantically rich foreground. Pooling-based approaches extract low-dimensional features but lack direct supervision on individual elements, often yielding suboptimal results. Correspondence-based methods that employ bipartite matching suffer from inherent instability, as supervisory signals vary across training iterations, leading to inefficient convergence.

Image tokenization has evolved significantly to address compression challenges. Variational Autoencoders (VAEs) pioneered mapping images into low-dimensional continuous latent distributions. VQVAE and VQGAN advanced this by projecting images into discrete token sequences, while VQVAE-2, RQVAE, and MoVQ introduced hierarchical latent representations through residual quantization. FSQ, SimVQ, and VQGAN-LC tackled representation collapse when scaling codebook sizes. Other methods like set modeling have evolved from traditional Bag-of-Words (BoW) representations to more complex techniques. Techniques like DSPN use Chamfer loss, while TSPN and DETR employ Hungarian matching, though these processes often generate inconsistent training signals.

Researchers from the University of Science and Technology of China and Tencent Hunyuan Research have proposed a fundamentally new paradigm for image generation through set-based tokenization and distribution modeling. Their TokenSet approach dynamically allocates coding capacity based on regional semantic complexity. This unordered token set representation enhances global context aggregation and improves robustness against local perturbations. Moreover, they introduced Fixed-Sum Discrete Diffusion (FSDD), the first framework to simultaneously handle discrete values, fixed sequence length, and summation invariance, enabling effective set distribution modeling. Experiments show the method’s superiority in semantic-aware representation and generation quality.

Experiments are conducted on the ImageNet dataset using 256 × 256 resolution images, with results reported on the 50,000-image validation set using the Frechet Inception Distance (FID) metric. TiTok’s strategy is followed for tokenizer training, applying data augmentations including random cropping and horizontal flipping. The model is trained on ImageNet for 1000k steps with a batch size of 256, equivalent to 200 epochs. Training incorporates a learning rate warm-up phase followed by cosine decay, gradient clipping at 1.0, and an Exponential Moving Average with a 0.999 decay rate. A discriminator loss is included to enhance quality and stabilize training, with only the decoder trained during the final 500k steps. MaskGIT’s proxy code facilitates the training process.

The results show key strengths of the TokenSet approach. Permutation-invariance is confirmed through both visual and quantitative evaluation. All reconstructed images appear visually identical regardless of token order, with consistent quantitative results across different permutations. This validates that the network successfully learns permutation invariance even when trained on only a subset of possible permutations. Each token integrates global contextual information with a theoretical receptive field encompassing the entire feature space by decoupling inter-token positional relationships and eliminating sequence-induced spatial biases. Moreover, the FSDD approach uniquely satisfies all desired properties simultaneously, resulting in superior performance metrics.

In conclusion, the TokenSet framework represents a paradigm shift in visual representation by moving away from serialized tokens toward a set-based approach that dynamically allocates representational capacity based on semantic complexity. A bijective mapping is established between unordered token sets and structured integer sequences through a dual transformation mechanism, allowing effective modeling of set distributions using FSDD. Moreover, the set-based tokenization approach offers distinct advantages, introducing possibilities for image representation and generation. This direction opens new perspectives for developing next-generation generative models, with future work planned to analyze and unlock the full potential of this representation and modeling approach.


Check out the Paper and GitHub Page. All credit for this research goes to the researchers of this project. Also, feel free to follow us on Twitter and don’t forget to join our 85k+ ML SubReddit.

The post TokenSet: A Dynamic Set-Based Framework for Semantic-Aware Visual Representation appeared first on MarkTechPost.

Share This Article
Twitter Email Copy Link Print
Previous Article GHA Discovery D$100 Bonus For Staying At Minor Hotels Brands Twice March 26 – September 30, 2025 (Book By May 20) GHA Discovery D$100 Bonus For Staying At Minor Hotels Brands Twice March 26 – September 30, 2025 (Book By May 20)
Next Article United & Chase Refresh Cobranded Card Portfolio With Higher fees & Structured Credits United & Chase Refresh Cobranded Card Portfolio With Higher fees & Structured Credits
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Your Trusted Source for Accurate and Timely Updates!

Our commitment to accuracy, impartiality, and delivering breaking news as it happens has earned us the trust of a vast audience. Using RSS feeds, we aggregate news from trusted sources to ensure real-time updates on the latest events and trends. Stay ahead with timely, curated information designed to keep you informed and engaged.
TwitterFollow
TelegramFollow
LinkedInFollow
- Advertisement -
Ad imageAd image

You Might Also Like

FHA cites AI emergence as it ‘archives’ inactive policy documents

By capernaum

Better leans on AI, sees first profitable month since 2022

By capernaum
A Step-by-Step Guide to Deploy a Fully Integrated Firecrawl-Powered MCP Server on Claude Desktop with Smithery and VeryaX
AI

A Step-by-Step Guide to Deploy a Fully Integrated Firecrawl-Powered MCP Server on Claude Desktop with Smithery and VeryaX

By capernaum
Implementing an LLM Agent with Tool Access Using MCP-Use
AI

Implementing an LLM Agent with Tool Access Using MCP-Use

By capernaum
Capernaum
Facebook Twitter Youtube Rss Medium

Capernaum :  Your instant connection to breaking news & stories . Stay informed with real-time coverage across  AI ,Data Science , Finance, Fashion , Travel, Health. Your trusted source for 24/7 insights and updates.

© Capernaum 2024. All Rights Reserved.

CapernaumCapernaum
Welcome Back!

Sign in to your account

Lost your password?