String Operations & Slicing Mechanics
Master string immutability, slicing edge cases, memory interning, formatting traps, and method evaluation quirks.
Theoretical Mechanics & Memory Architecture
Strings in Python (PyUnicodeObject) are immutable sequences of Unicode code points governed by PEP 393's Flexible String Representation. CPython inspects the maximum code point ordinal during creation to allocate the smallest possible character width: 1 byte (Latin-1/ASCII), 2 bytes (UCS-2), or 4 bytes (UCS-4). Understanding immutability, string interning, and memory allocation prevents quadratic scaling bottlenecks in algorithmic problem solving.
CPython maintains an internal hash table of interned strings. Short ASCII literals, dictionary keys, and valid identifiers are automatically interned, allowing fast pointer comparison (is) rather than byte-by-byte traversal (==). Because string objects are immutable, any transformation requires allocating a brand-new PyUnicodeObject and copying character bytes.
Repeatedly concatenating strings inside a loop with s += char allocates a new string buffer on every single iteration, copying all previous characters and resulting in O(N^2) quadratic time complexity. The idiomatic, linear-time approach is accumulating segments in a list and invoking ''.join(segments), which calculates total required memory up front and executes in a single pass.
Interactive Dry-Run Execution Workspace
Slicing Out of Bounds vs Indexing Out of Bounds
Trace CPython execution step-by-step and deduce the exact stdout string emitted by this snippet.
s = "abcde"
print(s[2:100], s[-100:2])Frequently Asked Questions on String Operations & Slicing Mechanics
Why is ''.join(list) significantly faster than repeated += string concatenation?
Repeated += concatenation reallocates a new buffer and copies accumulated characters on every single loop iteration (O(N^2)). ''.join() calculates total required length across all strings in advance, allocates one buffer of exact size, and copies bytes in a single O(N) pass.
How does PEP 393 optimize string memory in Python 3?
Instead of storing every character as 4-byte UTF-32, PEP 393 checks the maximum character ordinal in the string: strings with characters up to 255 use 1 byte per character (Latin-1); up to 65535 use 2 bytes (UCS-2); and emoji or rare scripts use 4 bytes (UCS-4).
When does CPython intern strings, and why is using 'is' for string equality discouraged?
CPython automatically interns identifiers, dictionary keys, and compile-time constants. Dynamically constructed strings (e.g. from user input or file reads) are not interned, so 'a is b' can evaluate to False even when 'a == b' is True. Always use '==' for value equality.