#12586: add provisional email policy with new header parsing and folding.

When the new policies are used (and only when the new policies are explicitly used) headers turn into objects that have attributes based on their parsed values, and can be set using objects that encapsulate the values, as well as set directly from unicode strings. The folding algorithm then takes care of encoding unicode where needed, and folding according to the highest level syntactic objects. With this patch only date and time headers are parsed as anything other than unstructured, but that is all the helper methods in the existing API handle. I do plan to add more parsers, and complete the set specified in the RFC before the package becomes stable.
2025-11-24 20:30:18 +00:00 · 2012-05-25 18:42:14 -04:00 · 2012-05-25 18:42:14 -04:00 · 0b6f6c82b5
commit 0b6f6c82b5
parent 0fa2edd08f
16 changed files with 6992 additions and 114 deletions
--- a/Lib/email/utils.py
+++ b/Lib/email/utils.py
@ -62,6 +62,13 @@ escapesre = re.compile(r'[\\"]')
 _has_surrogates = re.compile(
    '([^\ud800-\udbff]|\A)[\udc00-\udfff]([^\udc00-\udfff]|\Z)').search

+# How to deal with a string containing bytes before handing it to the
+# application through the 'normal' interface.
+def _sanitize(string):
+    # Turn any escaped bytes into unicode 'unknown' char.
+    original_bytes = string.encode('ascii', 'surrogateescape')
+    return original_bytes.decode('ascii', 'replace')
+

 # Helpers