Skip to content

Unquote and unescape Content-Type: charset [RFC7230] #267

Description

@forthrin
uri = URI('http://127.0.0.1:8080')
http = Net::HTTP.new(uri.hostname, uri.port)
http.response_body_encoding = true
http.send_request('GET', uri.request_uri)
/opt/homebrew/lib/ruby/gems/4.0.0/gems/net-http-0.9.1/lib/net/http/response.rb:381:in 'String#force_encoding': unknown encoding name - "UTF-8" (ArgumentError)
    @body.force_encoding(enc) if enc
	from /opt/homebrew/lib/ruby/gems/4.0.0/gems/net-http-0.9.1/lib/net/http/response.rb:381:in 'Net::HTTPResponse#read_body'
$ curl -I http://127.0.0.1:8080/
Content-Type: text/html; charset="UTF-8"
Server: WEBrick/1.9.2 (Ruby/4.0.1/2026-01-13)

$ open https://www.rfc-editor.org/rfc/rfc7230#section-3.2.6

[3.2.6](https://www.rfc-editor.org/rfc/rfc7230#section-3.2.6). Field Value Components
quoted-string = DQUOTE *( qdtext / quoted-pair ) DQUOTE

$ git diff
diff --git a/lib/net/http/header.rb b/lib/net/http/header.rb
index 797a3be..fe026ae 100644
--- a/lib/net/http/header.rb
+++ b/lib/net/http/header.rb
@@ -762,3 +762,3 @@ module Net::HTTPHeader
       k, v = *param.split('=', 2)
-      result[k.strip] = v.strip
+      result[k.strip] = v.strip.gsub(/\A"(.*?)(?<!\\)"\z/) { $1.gsub(/\\(.)/, '\1') }.strip
     end
http.send_request('GET', uri.request_uri)
# #<Net::HTTPOK 200 OK readbody=true>

Activity

  1. jeremyevans commented on Jan 16, 2026

    @jeremyevans
    Contributor

    I agree that this is an issue we should fix.

  2. forthrin commented on Jul 11, 2026

    @forthrin
    Author

    @jeremyevans: Been six months. Isn't this quite a straight forward fix?

  3. jeremyevans commented on Jul 11, 2026

    @jeremyevans
    Contributor

    The fix is fairly straight forward. Note that the the existing code does not correctly handle the case where there are multiple parameters and = or appears inside a quoted value of one of the parameters. Fixing that correctly requires writing an actual parser. Granted, content-type should only contain at most two parameters (charset and boundary), and in the net/http context, I think only charset makes sense (boundary is used for form submission request bodies, not generally for responses). I don't think any standard charset uses =, so it's probably safe. I'll submit a pull request for this, but realize that I am not a maintainer, so you'll have to wait for a maintainer to review/merge.

  4. forthrin commented on Jul 11, 2026

    @forthrin
    Author

    Thanks for jumping in. When a core component with few open issues called by about any Ruby HTTP request has an exception that has a simple reason and a simple solution, a merge would seem imminent.

  5. duerst commented on Jul 12, 2026

    @duerst
    Member

    Granted, content-type should only contain at most two parameters (charset and boundary), and in the net/http context, I think only charset makes sense (boundary is used for form submission request bodies, not generally for responses). I don't think any standard charset uses =, so it's probably safe.

    = is indeed not allowed in charset names. See RFC 2978, Section 2.3 (https://datatracker.ietf.org/doc/html/rfc2978#section-2.3). Also, of course, none of the registered charsets at https://www.iana.org/assignments/character-sets/character-sets.xhtml contains a =.

  6. forthrin commented on Sep 15, 2026

    @forthrin
    Author

    @duerst: Any progress here? Seems like quite a simple fix for fundamental compliance.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions