Security & devFreeNo signup
Free Base64 encoder and decoder for text. Supports standard and URL-safe alphabets. Runs locally in your browser.
Last updated 3 October 2026
A transport encoding, not a protection. Everything this page does is reversible by anyone, with no key and no password, including by this page. RFC 4648 section 12: base encoding “does not provide any computational confidentiality”. The quotations below were read from RFC 4648 and from the WHATWG and W3C specifications that define the browser functions actually doing the work, on 2 October 2026.
Put QQ== in the box above and decode it. You get A. Now try QR==, then QS==, then Qf==. Every one of them decodes to the same single letter, and every one of them is accepted without complaint.
All sixteen strings that this page decodes to the one byte 0x41
This is neither a bug in the page nor an accident of the browser. Both of the specifications involved describe it, from opposite directions.
The cause is arithmetic. Base64 moves three bytes at a time into four characters of six bits each. When the input does not divide by three the last group is short: one leftover byte supplies 8 bits into a space of 12, two leftover bytes supply 16 into a space of 18, and the spare bits carry nothing. RFC 4648 section 3.5 tells encoders what to do with them, “These pad bits MUST be set to zero by conforming encoders”, then says what follows when that does not hold: “there is no canonical representation of base-encoded data, and multiple base-encoded strings can be decoded to the same binary data.”
So a conforming encoder produces only QQ==. Decoders are the loose end. RFC 4648 leaves it as a choice: “decoders MAY chose to reject an encoding if the pad bits have not been set to zero.” The decoder behind this page is atob, and the WHATWG Infra Standard, which specifies it, chose not to reject. Its algorithm discards the spare bits, and the standard flags the consequence in its own text: “The discarded bits mean that, for instance, ‘YQ’ and ‘YR’ both return a.” The same standard is candid that this is a deliberate divergence, describing the algorithm as “different from the RFC as it defines error handling for certain inputs”.
Why it matters outside a curiosity: RFC 4648 section 12 lists it as a security consideration. “When padding is used, there are some non-significant bits that warrant security concerns, as they may be abused to leak information or used to bypass string equality comparisons or to trigger implementation problems.” If any code anywhere compares Base64 strings to decide something, sixteen strings pass for one byte. Compare the decoded bytes.
This is the misconception the format exists inside, and RFC 4648 section 12 addresses it directly, with a scenario rather than a scolding:
RFC 4648, section 12, Security Considerations
The scenario is not hypothetical. HTTP Basic authentication sends a username and password as one Base64 string, so YWRtaW46aHVudGVyMg== is a complete credential, and it is the string admin:hunter2. Decode it above. A bug report, a support ticket or a pasted log line containing that header has disclosed the password, and it does not look like it has, which is the whole problem. The encoding is doing exactly its job: it made bytes safe to transport and nothing else.
The third quotation is the subtle one: Base64 does not merely fail to protect, it slightly helps an attacker, by making the plaintext longer and giving it a recognisable statistical fingerprint. What needs to be confidential needs encryption, and the encrypted bytes can then be Base64 encoded for transport.
The Alphabet selector above is not a formatting preference. It chooses between two tables that RFC 4648 defines separately, and the difference between them is exactly two entries out of sixty-four.
RFC 4648 Table 1 against Table 2 · the only two rows that differ
The reason for the second table is that + and / are hostile in two specific places. RFC 4648 section 3.4 names both: “the non-alphanumeric characters (in particular, ‘/’) may be problematic in file names and URLs”, and “certain characters, notably ‘+’ and ‘/’ in the base 64 alphabet, are treated as word-breaks by legacy text search/index tools”. A slash in a path is a directory separator, and a plus in a query string has historically meant a space. Section 5 records the alternatives dropped on similar grounds, ~ and ., both awkward in filenames.
The naming is a real request, not pedantry. RFC 4648 section 5: “This encoding may be referred to as ‘base64url’. This encoding should not be regarded as the same as the ‘base64’ encoding and should not be referred to as only ‘base64’.” An earlier version of this page called it “URL-safe Base64”, which is the phrasing that sentence asks people to stop using. The proper name is base64url, and asking a colleague for “the base64” of something is genuinely ambiguous.
Each setting rejects the other's two characters, and the error says which alphabet the stray one came from. That is what RFC 4648 section 3.3 asks for: “implementations MUST reject the encoded data if it contains characters outside the base alphabet”. So a - or a _ under Standard, and a + or a / under URL-safe, are refused with a message naming the character, the alphabet it is absent from, and the setting that would accept it.
Until 3 October 2026 the URL-safe setting accepted + and / anyway, which made it quietly the more forgiving of the two. A string mixing the alphabets decoded without complaint into bytes that were not the ones it encoded, which is the worst available outcome: no error, wrong answer. If a paste will not decode under one setting, the message now tells you whether the other one is the answer, so there is no longer any reason to try both and guess.
If the alphabets differ in two of sixty-four characters, how often does the choice actually change anything? The answer is lopsided in a way that explains why this bug survives in production code for years, and it is measurable.
Does switching alphabet change the output? · measured 2 October 2026
Thirty-two bytes is not arbitrary: it is the length of an HMAC-SHA256 output, the signature on the most common kind of JWT, and of a great many API keys. Those encode to 43 base64url characters, about three in four containing a - or a _.
ASCII JSON almost never does. The byte patterns in {“sub”:…} keep nearly every six-bit group inside the fifty-two letters and ten digits the two tables share, so the header and payload of a JWT usually look character-for-character identical under either alphabet. Nothing in 20,000 realistic claim sets differed.
The practical shape of the bug follows from those two numbers. Code that uses the wrong alphabet on a JWT works perfectly against the readable parts, in testing and in review, and fails on the signature, which is the one segment nobody reads by eye. The failure then presents as “the signature does not verify”, which sends everyone looking at keys and clocks rather than at a character table.
The trailing = confuses people because it looks like a terminator and is not. It is a length marker, and RFC 4648 section 4 permits exactly three cases, because “a full encoding quantum is always completed at the end of a quantity”.
The three final-quantum cases of RFC 4648 section 4
It follows that a Base64 string's length is always a multiple of four, that it never ends in three or more =, and that a length leaving a remainder of one when divided by four is impossible. If you are writing a validator, those three facts are most of it.
Whether the padding must be there depends on who is asking. RFC 4648 section 3.2 sets the default: “Implementations MUST include appropriate pad characters at the end of encoded data unless the specification referring to this document explicitly states otherwise”, because “when assumptions about the size of transported data cannot be made, padding is required to yield correct decoded data.” Section 5 gives the standing exception for URLs: “the pad character ‘=’ is typically percent-encoded when used in an URI, but if the data length is known implicitly, this can be avoided by skipping the padding.”
That is why the URL-safe mode above strips the padding and the Standard mode keeps it. It is also why JWT segments never carry it. Decoding here is more relaxed than encoding in one specific way: a string carrying no padding at all is padded back out before decoding, so aGVsbG8 and aGVsbG8= both give hello. Padding that is present but wrong is a different case and is refused, because a string padded to a length Base64 cannot have is damaged rather than merely abbreviated.
A string with the wrong number of =, such as aGk==, is refused, and the message names the padding rather than the characters. Every character in aGk== is in the alphabet; what is wrong is that five characters cannot be a padded Base64 string, since padding exists to fill a multiple of four. Until 3 October 2026 this page answered “Invalid Base64 characters” here, which sent people looking at the wrong half of their data.
The expansion matters if you are putting bytes inside something with a size limit. Three bytes become four characters, and the final group is always completed whether or not it is full.
Measured output length · RFC 4648 section 4 arithmetic, 24 bits in to 4 characters out
That ratio is where the mail-limit folklore comes from: a service capping a message at 25 MB is capping the encoded form, so the attachable file is nearer 18 MB. The same arithmetic applies to an image inlined as a data URL, a certificate in PEM form, and anything stored Base64 encoded in a database column.
One related constraint, since this page emits a single unbroken string. RFC 4648 section 3.1: “Implementations MUST NOT add line feeds to base-encoded data unless the specification referring to this document explicitly directs base encoders to add line feeds after a specific number of characters.” The 76-character lines in mail bodies and the 64-character lines in PEM files are requirements of MIME and of PEM, not of Base64. Unbroken is correct for a data URL, a token or an HTTP header, and needs wrapping only if you are hand-assembling a MIME part.
Base64 is defined over bytes. Until 3 October 2026 this page was built on text, and that mismatch cost data in two separate ways while reporting success in both. It now decodes to bytes and treats text as one view of them. The output box shows text only when the bytes genuinely are text and the box can hold them unchanged; otherwise it shows the bytes as hex and the status line says which of those two things failed.
Decoding the first bytes of a PNG file · measured on this page in a browser, 3 October 2026
Those are the real first sixteen bytes of every PNG file, and the round trip returns exactly what went in. For comparison, the same paste on the code this page carried until today returned ef bf bd 50 4e 47 0a 1a 0a 00 00 00 0a 49 48 44 52 and Decoded · 15 chars in the success colour, and pressing the same button gave back 77+9UE5HChoKAAAACklIRFI=, a different string from the one typed in.
The first byte of a PNG is 0x89, which is not a legal start to a UTF-8 sequence. The WHATWG Encoding Standard gives a decoder two error modes for that situation, “replacement” or “fatal”, and under “replacement” the instruction is to “push U+FFFD (�) to output”. The fatal option defaults to false, so accepting the default meant 0x89 silently became U+FFFD, which re-encodes as the three bytes EF BF BD. This page now sets fatal to true and asks the question before answering: if the bytes are not valid UTF-8, it says so rather than handing you a transcription of them in green.
The second loss had nothing to do with binary data and happened to perfectly ordinary text. The two boxes above are HTML textarea elements, and a textarea normalises line endings: a carriage return put into one comes back out as a line feed, or disappears. That is what the HTML standard requires rather than a browser quirk, so no amount of care inside the page can put a carriage return into the box. Two consequences follow, and both are now visible rather than silent. On decode, bytes containing a carriage return are shown as hex. On encode, a Line endings control asks for CRLF explicitly, and an Input is control accepts hex bytes, which skips the text box's normalisation altogether.
Carriage returns through the two boxes · measured in a browser, 3 October 2026
Both line-ending answers are correct for some input, which is why the control exists and is labelled rather than guessed at. The code this page carried until today offered only the second one, and produced YQpi for a PEM block while calling it a success. Anything downstream verifying a hash or a signature over those bytes failed with no hint as to why, which is the failure this page existed to prevent.
So the three rules this section used to give are retired, and one holds. A PEM certificate, a captured HTTP exchange or a block of mail headers can be encoded here faithfully by setting Line endings to CRLF, and any byte sequence at all can be encoded by setting Input is to hex. The one remaining limit is real: a browser page cannot read a file off your disk, so for a whole certificate or an image a command-line tool is still the better instrument, and for anything over a few kilobytes it is also the faster one. The difference is that this page no longer gives you a wrong answer in the success colour.
One counting note, while on the subject. The status line reports bytes, which is the unit Base64 operates on. It used to report the length of the decoded JavaScript string, which counts UTF-16 code units: one emoji counted as 2, and an emoji assembled from several code points joined by zero-width joiners counted several more. Measured now, the four-person family emoji reports 25 bytes, which is what it occupies in UTF-8, and the single smiling face reports 4 bytes.
A few controls and a text box. Everything below is outside what they can establish.
What the bytes were. It can now tell you whether they are valid UTF-8, which is a narrower question than whether they are text and a much narrower one than what they are. A PNG is not valid UTF-8 and is reported as such. A key in PEM form is text, so it decodes as text. A compressed blob that happens to be valid UTF-8 by coincidence will be shown as text with no warning, because by that test it is text.
Whether your line endings were CRLF to begin with. They survive now, in both directions, but only because you say which they were. The box cannot hold a carriage return, so it cannot infer that the file you copied from had one: the Line endings control is you telling the page, and it has no way to check you are right.
Which alphabet a token was written in. You choose that. Since ASCII JSON encodes identically under both, the page cannot infer it from the input, and a JWT's header will decode correctly under the wrong setting while its signature will not.
Whether the padding was originally there. Decoding pads a short string back out, so aGVsbG8 and aGVsbG8= are treated alike. If a protocol you are debugging cares about the exact bytes on the wire, this page has already normalised them.
Whether the encoding was canonical. Sixteen strings decode to the letter A and all sixteen are accepted. The page cannot tell you that the one you were given had non-zero pad bits, which is sometimes the signal that something upstream is generating Base64 by hand.
What the decoded content means, or whether it is safe. RFC 4648 section 3.3 notes that non-alphabet characters inside encoded data “may be exploited as a ‘covert channel’”. Decoding something does not make it safe to run, open or trust.
Whether what you decoded is still a live secret. A token or credential pasted here is now in your clipboard and this browser session, which is true of any tool and is the reason to rotate rather than to worry.
Nothing you type above leaves your browser. The work is done by btoa and atob, two functions built into the browser itself, in a few lines of JavaScript on this page; there is no server call, no upload and no logging of the contents of either box.
Every quotation above was read from the IETF or WHATWG document named, on 2 October 2026, rather than from a summary. The sixteen strings, the length table, the alphabet-divergence percentages, the carriage-return table and the PNG round trip were measured by running this page's own encode and decode functions in a browser. Part of the QuikUtil tools collection; the JWT Decoder applies base64url to a token's three segments, and URL Encode / Decode covers percent-encoding, a different thing often confused with this one.
Carrying bytes through something that will only accept text. RFC 4648 describes it as designed “to represent arbitrary sequences of octets in a form that allows the use of both upper- and lowercase letters but that need not be human readable”, using “a 65-character subset of US-ASCII”. Hence mail attachments, data URLs, JWT segments, Basic credentials and certificate files. It is a transport wrapper, nothing more.
No, and RFC 4648 section 12 is blunt about the consequence: “Base encoding visually hides otherwise easily recognized information, such as passwords, but does not provide any computational confidentiality. This has been known to cause security incidents when, e.g., a user reports details of a network protocol exchange … and accidentally reveals the password because she is unaware that the base encoding does not protect the password.” The same section adds that base encoding “adds no entropy to the plaintext”. YWRtaW46aHVudGVyMg== is the string admin:hunter2, and anyone can see that with this page.
Exactly two characters out of sixty-four. RFC 4648 Table 1 gives index 62 as + and index 63 as /; Table 2, the URL and filename safe alphabet, gives the same indices as - (minus) and _ (underline). Everything else, all fifty-two letters and ten digits, is identical. The specification also asks you not to use the shorter name for it: “This encoding may be referred to as ‘base64url’. This encoding should not be regarded as the same as the ‘base64’ encoding and should not be referred to as only ‘base64’.”
Look for a +, a /, a - or a _, and accept that most of the time there will not be one. Measured while writing this page: of 100,000 random 32-byte values, which is the size of an HMAC-SHA256 signature or a typical API key, 73.5% encoded differently under the two alphabets. Of 20,000 ASCII JSON claim sets of the shape a JWT carries, not one differed. So the readable half of a JWT looks the same either way and the signature usually does not, which is why this mistake hides for so long and then surfaces in the part you cannot read.
Because the encoder works in 24-bit groups and the last group is often short. RFC 4648 section 4 allows exactly three endings: an input that is a multiple of three bytes ends with no padding, an input leaving two bytes ends with one =, and an input leaving one byte ends with two =. So abc becomes YWJj, ab becomes YWI= and a becomes YQ==. Padding is normally required (“Implementations MUST include appropriate pad characters at the end of encoded data unless the specification referring to this document explicitly states otherwise”), and section 5 gives the common exception: “the pad character ‘=’ is typically percent-encoded when used in an URI … but if the data length is known implicitly, this can be avoided by skipping the padding”. The URL-safe mode on this page drops it for that reason, which is also what JWT requires.
Yes, and that is the most surprising thing about the format. Sixteen distinct strings decode to the single byte that spells the letter A: QQ== through Qf==. Four decode to hi: aGk=, aGl=, aGm= and aGn=. The cause is in RFC 4648 section 3.5, which requires encoders to set the leftover pad bits to zero but notes that otherwise “multiple base-encoded strings can be decoded to the same binary data”. The browser decoder behind this page is specified to ignore those bits rather than reject them, and the WHATWG Infra Standard names the effect itself: “the discarded bits mean that, for instance, ‘YQ’ and ‘YR’ both return a”. If you compare Base64 strings for equality anywhere that matters, compare the decoded bytes instead.
Characters outside the alphabet are rejected, as RFC 4648 section 3.3 requires: paste !!! or a - while the Standard alphabet is selected and you will get an error. But three things get through. Non-canonical pad bits, as above. ASCII whitespace anywhere inside the string, which the browser decoder removes before it starts. And valid Base64 whose bytes are not valid UTF-8 text, which is the one to watch, because it is reported as a success.
A third larger. Four output characters carry three input bytes, so the overhead settles at 33.33%: 1,024 bytes becomes 1,368 characters and five megabytes becomes 6,990,508. Small inputs are worse, because the final group is always completed, so one byte becomes four characters. That ratio is why a 25 MB mail limit means roughly 18 MB of actual file.
No. The conversion is done by btoa and atob, two functions built into your browser, in a few lines of JavaScript on this page. Nothing is sent to a server and nothing is stored. Be aware of the other direction, though: a password or token you paste into any page is then in your clipboard and your browser session. If you decode a live credential here, treat it as one that has been handled and rotate it.