Fix Garbled HTML in DEV.to Comments for Drafting Replies with Python
Job to be done: Process raw HTML comments from an API into plain text for drafting replies
🇳🇬 Ways to use this in Nigeria
Ideas to get you started, adapt to your situation.
- Student
Clean garbled HTML from DEV.to comments to draft replies to your technical articles.
- 9-5 employee
Process raw HTML comments from DEV.to API into plain text for drafting replies to your articles.
What you’ll get
You will get a corrected Python script that takes raw HTML comments from the DEV.to API and converts them into clean, readable plain text. This ensures that special characters like ampersands and angle brackets are displayed correctly, making it easier to draft replies to comments on your DEV.to articles.
This approach works because it uses Python’s built-in html module to correctly decode HTML entities that the DEV.to API sometimes returns in a garbled format.
Tools you need
- Python (free): A programming language used to write the script.
- DEV.to API (free): Provides the raw comment data from your DEV.to articles.
Steps
-
Set up your Python environment: If you don’t have Python installed, download and install it from the official website. You can run Python scripts on your computer or on your phone using apps like Pydroid 3 for Android.
-
Create a new Python file: Open a text editor or your Python IDE and create a new file named
reply_comments.py. -
Add the corrected
strip_htmlfunction: Copy and paste the following Python code into yourreply_comments.pyfile. This code includes the necessary import and the updatedstrip_htmlfunction that correctly handles HTML entities.import re import html def strip_html(h): """Plain text of a dev.to comment's body_html, for drafting replies. dev.to's rendered HTML entity-escapes a commenter's own literal <, >, and " characters right alongside the actual tags it wraps the comment in. Stripping tags alone leaves those entities untouched, so a comment quoting code with generics (List <String>) or using " " came back as literal " < " / " > " text in the exact field a reply gets drafted from. """ # First, remove HTML tags using a regular expression no_tags = re.sub(r'<[^>]+>', ' ', h) # Then, decode HTML entities like &, <, >, ' unescaped = html.unescape(no_tags) # Finally, collapse multiple whitespace characters into a single space and strip leading/trailing whitespace return re.sub(r'\s+', ' ', unescaped).strip() # Example usage (optional, for testing): # raw_html_comment = "<p>Isn't it faster with a Q&A cache? Try List <String> instead.</p>" # plain_text_comment = strip_html(raw_html_comment) # print(plain_text_comment) # Expected output: Isn't it faster with a Q&A cache? Try List <String> instead.You should see the Python code pasted into your file. If you run the example usage at the bottom, you should see the corrected plain text output.
-
Integrate with your script: If you have an existing script that pulls comments from the DEV.to API, replace your old
strip_htmlfunction with this new one. Ensure your script calls this function on thebody_htmlfield it receives from the API.
Original source
This workflow is based on a post by enjoy_kumawat on DEV Community. The author identified an issue where their script was not correctly displaying special characters in comments pulled from the DEV.to API, leading to garbled text when drafting replies.
Notes & variations
- Free tier viability: This workflow is entirely free, using Python and the DEV.to API, which has no usage limits for this purpose.
- Common mistake: Forgetting to import the
htmlmodule. Without it,html.unescapewill not be recognized, and the entities will not be decoded. - Tip for better results: If you encounter other unusual characters or encoding issues, you might need to explore more advanced text cleaning libraries in Python, but for most common HTML entities, this corrected function should suffice.