MEGA Hub

TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters

Authors

Do you know Honglie Wang?You can claim authorship or link another user.Do you know Jia Sun?You can claim authorship or link another user.Do you know Zijun Li?You can claim authorship or link another user.Do you know Junlong Wu?You can claim authorship or link another user.Do you know Pengcheng Wei?You can claim authorship or link another user.Do you know Jiyuan Wang?You can claim authorship or link another user.Do you know Yongrui Heng?You can claim authorship or link another user.Do you know Boheng Zhang?You can claim authorship or link another user.Do you know Huaiqing Wang?You can claim authorship or link another user.Do you know Dewen Fan?You can claim authorship or link another user.Do you know Qianqian Gan?You can claim authorship or link another user.Do you know Fan Yang?You can claim authorship or link another user.Do you know Tingting Gao?You can claim authorship or link another user.Do you know Yan-Ming Zhang?You can claim authorship or link another user.

Abstract

Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition. Despite recent progress in instruction-based image editing, general-purpose models remain unreliable in this setting: they often omit or incorrectly render the target text, place it over salient products or pre-existing content, and produce structurally distorted or visually inconsistent glyphs. We introduce \textbf{TextRefine}, a task-aligned post-training framework that combines supervised fine-tuning with operation-specific reward optimization to address these complementary failure modes. For text insertion, our text-span-level reward jointly assesses semantic fidelity and target-span coverage, penalizes spatial conflicts with products and existing text, and employs a gated structural constraint to preserve non-text regions. For text replacement, our glyph-level reward leverages the connectionist temporal classification (CTC) posterior of the target character to provide graded supervision for fine-grained defects, including missing strokes, structural deformations, and confusion among visually similar characters. We further introduce \textbf{OpenTextEdit}, a dataset comprising 100K images for text editing in product posters, with multi-text layouts, detailed text attributes, product masks, and challenging low-frequency characters. Extensive experiments on both insertion and replacement demonstrate that TextRefine consistently outperforms the evaluated image editing baselines in textual fidelity, placement reliability, and glyph quality while better preserving source-image content.

Community

00