Converting Telegram Messages to X Posts: Weighted Character Counts, Entities, and URLs in 2026
Operators of Telegram→X schedulers often run into the same failure. A draft looks short in Telegram, gets approved, and then X rejects it at publish time, or it goes out with a dead link because the clickable text in Telegram had no visible URL. The cause is almost always text conversion. Telegram and X count characters differently, carry links differently, and support different formatting. This guide explains how to turn a Telegram message, meaning its text plus its entities, into valid X Post text, and how to show the operator an accurate count before approval. It builds on our posts about approval gates and duplicate-content fingerprints, and it is engineering guidance, not a reach or growth promise.
Two platforms, two ways of counting
Start with what each side documents. Re-check these pages before you ship, because both APIs change:
- Telegram limits are counted after entity parsing. The sendMessage reference allows message text of 1–4096 characters “after entities parsing,” and media captions allow 0–1024. An operator can paste a 3,000-character draft into your bot without Telegram complaining.
- Telegram entity positions are UTF-16 offsets. In the MessageEntity object,
offsetandlengthare measured “in UTF-16 code units.” Entity types includeurl,text_link(clickable text with a hidden URL),mention,hashtag,cashtag,bot_command,bold,italic,spoiler,code,pre,custom_emoji, and others. - X counts weighted characters. X’s Counting Characters page says Posts can contain up to 280 characters, but not every character counts the same. Latin letters, punctuation, and common symbols weigh 1. Emoji weigh 2 no matter how many code points or zero-width joiners make them up. CJK characters weigh 2, and other Unicode defaults to 2. Every valid URL is wrapped by t.co and counts as 23 characters, regardless of its real length. Text is measured after Unicode Normalization Form C (NFC).
- X points to a reference implementation. The same page recommends the open-source twitter-text library, whose
parseTweet/parse_tweetreturnsweightedLengthandvalid. The weight table is in its config/v3.json.
So a Telegram length check tells you almost nothing about whether X will accept the Post. You have to compute X’s count yourself, on the converted text.
What the v3 weight table actually means
The v3 config sets maxWeightedTweetLength to 280, scale to 100, defaultWeight to 200, and transformedURLLength to 23. Only four code-point ranges get weight 100, which means 1 visible character:
U+0000–U+10FF: Latin, Greek, Cyrillic, Armenian, Hebrew, Arabic, Devanagari and other Indic scripts, Thai, Georgian, and more.U+2000–U+200D: typographic spaces and zero-width characters.U+2010–U+201F: hyphens, en and em dashes, and curly quotes.U+2032–U+2037: primes.
Everything else weighs 2. That has consequences an operator won’t see in Telegram. The arrow → (U+2192) and the ellipsis … (U+2026) each count as 2. Hangul and CJK count as 2. The “𝐛𝐨𝐥𝐝” letters some tools generate to fake formatting come from the Mathematical Alphanumeric Symbols block, so they count as 2 each, and screen readers often read them poorly. A caption full of arrows and styled letters can be noticeably “longer” on X than it looks in Telegram.
Slice Telegram entities in UTF-16, not code points
The most common conversion bug is slicing entities with the wrong unit. JavaScript strings are UTF-16, so text.slice(offset, offset + length) lines up with Telegram’s offsets. Python strings index by code point, so any emoji outside the Basic Multilingual Plane, which is two UTF-16 units, shifts every later entity by one. The result is a link wrapped around the wrong word, or a mention that loses its last letter. In Python, slice on the UTF-16 encoding:
def tg_slice(text: str, offset: int, length: int) -> str:
b = text.encode("utf-16-le")
return b[offset * 2:(offset + length) * 2].decode("utf-16-le")
Process entities from the end of the string toward the start when you rewrite text, so earlier offsets stay valid. Apply NFC only after you have finished applying entities, because normalization can change UTF-16 lengths.
Map each entity type deliberately
X Post text is plain text, so every Telegram entity needs an explicit rule. Here is a mapping we consider sensible for operator-run schedulers:
url: keep it. It counts as 23 on X whatever its length. Note that twitter-text also detects many domains without a scheme, so “example.com” in a caption may count as 23 too.text_link: the URL lives in the entity’surlfield, not in the text. If you copy only the text, the link disappears. Either append the URL (that costs 23 plus a space) or flag the draft so the operator chooses. Do not drop it silently.mention/text_mention: a Telegram@usernameis not an X handle. Copied as-is, it may tag an unrelated X account. Map known handles from a table you maintain, and strip the@from anything else.hashtag/cashtag: keep them, but remove Telegram’s@chatusernamesuffix forms such as#tag@channel.bold,italic,underline,strikethrough,code,pre: strip them to plain text. Don’t substitute styled Unicode letters.spoiler: X has no equivalent, so the hidden text would be fully visible. Block approval until the operator rewrites it.bot_command: remove the leading/scheduleor/postcommand from the text you publish.custom_emoji: the message text already holds a standard fallback emoji, which weighs 2 on X. Keep that.
text_link (lost URLs), mention (wrong X account), and spoiler (exposed text).Count before approval, and again before publish
Put the conversion and the count in front of the human gate, not behind it:
- Convert on receipt. When a draft arrives, build the X text with the mapping above, normalize it to NFC, and run twitter-text’s parser on the result. Store the converted text and its
weightedLengthwith the draft. - Show the operator the result, not the input. Reply in Telegram with the converted text, the count (for example, “271/280”), and any warnings: an appended link, a stripped mention, a blocked spoiler. Approval then covers what will actually be published.
- Reject over-length drafts early. If
validis false, say by how much it is over and don’t offer the approve button. Do not auto-truncate. A Post cut mid-sentence or mid-URL is worse than a delayed one. - Re-validate in the worker. Just before calling Create Post, recompute the count on the exact payload. If an edit or a config change slipped in after approval, the job should fail closed into your dead-letter queue (see failed-publish retries and dead letters) rather than burn a request on a guaranteed rejection.
- Hash the converted text. Compute duplicate fingerprints and idempotency keys on the normalized X text, not the raw Telegram text. Otherwise two drafts that differ only in bold formatting count as “different.” See client-request-id idempotency.
Attached media counts as 0 characters, according to X’s page, so moving an image link into a real upload saves 23 characters. Our media upload pipeline post covers that path. Some account tiers allow longer Posts. Treat that as a per-account setting you look up, not a constant you hard-code.
Test cases worth keeping in CI
- An emoji before a
text_link, to catch Python code-point slicing. - A ZWJ family emoji (👨👩👧👦), which should count as 2.
- NFD “café” versus NFC “café”, which should both count 4.
- A 200-character URL, which should count 23.
- A caption with
→and…, where each should count 2. #tag@channel, which should become#tag.- A spoiler, which should block approval.
- Exactly 280 weighted characters, which should pass, and 281, which should fail.
Pin your twitter-text version and record it with each published job. That way, if counting rules change later, you can explain why an older Post was accepted.
Practical checklist
- Slice Telegram entities in UTF-16 units, working from the end of the string.
- Map every entity type explicitly. Never lose
text_linkURLs, never copy mentions across platforms blindly, and block spoilers. - Normalize to NFC, then count with twitter-text. Don’t write your own
len()check. - Show the operator the converted text and its weighted count before approval, and never auto-truncate.
- Re-validate in the worker, and send failures to the dead-letter queue.
- Fingerprint and dedupe on the converted text.
Key takeaways
- Telegram’s 4096-character limit and X’s 280 weighted characters measure different things. Only X’s count decides whether a Post is accepted.
- URLs always count 23, emoji and most non-Latin-1 symbols count 2, and text is measured after NFC normalization.
- Telegram entity offsets are UTF-16 code units. Code-point slicing silently corrupts links and mentions.
- Approval is only meaningful if the operator approves the exact converted text and count that the worker will publish.
Character weights, limits, and entity definitions summarized here come from X’s and Telegram’s public developer documentation and the twitter-text v3 configuration as of 2026-10-06. They can change, so confirm them on docs.x.com and core.telegram.org before relying on them. No affiliate links. Last verified 2026-10-06.