An incredibly awful security vulnerability just got revealed in MongoDB.
So much that it got named after HeartBleed.
MongoBleed is a vulnerability affecting all MongoDB versions from 2017 to... today.
The exploit is simple. It's a buffer over read bug due to compression. Here's how it works ๐
Clients can send compressed requests to MongoDB.
The client helpfully includes the uncompressed size of the message so the server knows exactly how much memory to allocate when decompressing.
The server allocates a memory buffer with the given space. Due to how memory management and garbage collection in programs work, this allocated memory may already contain sensitive information that was copied earlier and is considered garbage now (eg because it's unreferenced).
This is technically fine - every computer program works that way because it is assumed that whatever unclaimed memory exists there will be overwritten. Unfortunately thatโs exactly where the bug lies. ๐
The server stupidly trusts the clientโs provided uncompressed size. When a malicious client lies about the uncompressed size - e.g the actual decompressed size is 100 bytes, but the client says its 1MB - Mongo will treat the full 1MB block as the message.
It will unload the 100 byte decompressed msg into the buffer, yet treat the full 1MB block as the msg.
This is extremely problematic if you can get the server to return back parts of the 1MB block, because it could contain data you may not have access to.
That is exactly what the exploit does - it sends a badly-formatted BSON message. The server fails to parse it, and "helpfully" returns an error message containing the invalid message. The invalid message can be that whole 1MB block of foreign data.
To understand the exploit a bit better, you need to understand the MongoDB protocol.
โข Mongo also uses its own TCP wire format (i.e doesn't use HTTP, gRPC or the like).
โข BSON is Mongo's message format passed within the TCP wire format. BSON is basically JSON in binary form
โข Commands in Mongo don't have particular endpoints or RPC names - rather, they are simply JSON-like messages. The action is inferred from the first key of the JSON.
For example, an insert request looks like this:
`{ "insert": "users", "documents": [ { "name": "alice", "age": 30 } ] }`
Every request to the server is therefore decoded into the BSON format as itโs parsed.
Critically, BSON parsing of field names (which are strings) work by parsing the field until you hit a null terminator byte (0x00).
It works exactly like strings in C, which have their own rich history of vulnerabilities.
We can now tie things together:
1. The client lies to the the server that its request has a big uncompressed size, so the server allocates a large block of memory
2. The client sends an invalid BSON with a field which does NOT contain the null terminator (0x00)
3. The server naively tries to parse the BSON field in that allocated block until it hits the first null byte. The first null byte is encountered in some foreign data since the BSON literally doesn't have it
4. The server realizes this is a completely invalid BSON message so it responds with an error.
5. The error response contains the invalid BSON "field". Critically, the server parsed garbage data from the heap in step 3), so it returns that data in the response.
Congrats. If the garbage contains passwords or other sensitive info, youโve hacked MongoDB!
Hackers exploit this by sending many malicious requests per second and then attempting to reconstruct the pieces of garbage they received back.
Whatโs critical about this vulnerability is that it works on ANY internet-accessible unpatched instance of MongoDB. ๐
You don๏ฟฝ๏ฟฝ๏ฟฝt need to authenticate with the server, because this whole request/response parsing cycle happens before the server can even authenticate.
Obviously you canโt authenticate a malformed request which doesnโt contain credentials - so that path of the code never gets executed.
The server simply responds with an error response. It just so happens that this error response can contain sensitive data. ๐คทโโ๏ธ
Merry Christmas
๐ง๐ผ๐ฝ ๐ฎ๐ฌ ๐ฆ๐ค๏ฟฝ๏ฟฝ ๐พ๐๐ฒ๐ฟ๐ ๐ผ๐ฝ๐๐ถ๐บ๐ถ๐๐ฎ๐๐ถ๐ผ๐ป ๐๐ฒ๐ฐ๐ต๐ป๐ถ๐พ๐๐ฒ๐
Here is the list of the top 20 SQL query optimization techniques I found noteworthy:
1. Create an index on huge tables (>1.000.000) rows
2. Use EXIST() instead of COUNT() to find an element in the table
3. SELECT fields instead of using SELECT *
4. Avoid Subqueries in WHERE Clause
5. Avoid SELECT DISTINCT where possible
6. Use WHERE Clause instead of HAVING
7. Create joins with INNER JOIN (not WHERE)
8. Use LIMIT to sample query results
9. Use UNION ALL instead of UNION wherever possible
10. Use UNION where instead of WHERE ... or ... query.
11. Run your query during off-peak hours
12. Avoid using OR in join queries
14. Choose GROUP BY over window functions
15. Use derived and temporary tables
16. Drop the index before loading bulk data
16. Use materialized views instead of views
17. Avoid != or <> (not equal) operator
18. Minimize the number of subqueries
19. Use INNER join as little as possible when you can get the same output using LEFT/RIGHT join.
20. For retrieving the same dataset, frequently try to use temporary sources.
Do you know what is ๐ค๐๐ฒ๐ฟ๐ ๐ข๐ฝ๐๐ถ๐บ๐ถ๐๐ฒ๐ฟ? Its primary function is to determine ๐๐ต๐ฒ ๐บ๐ผ๐๐ ๐ฒ๐ณ๐ณ๐ถ๐ฐ๐ถ๐ฒ๐ป๐ ๐๐ฎ๐ to execute a given SQL query by finding the best execution plan. The query optimizer works by taking the SQL query as input and analyzing it to determine how best to execute it. The first step is to parse the SQL query and create a syntax tree. The optimizer then analyzes the syntax tree to determine how to run the query.
Next, the optimizer generates ๐ฎ๐น๐๐ฒ๐ฟ๐ป๐ฎ๐๐ถ๐๐ฒ ๐ฒ๐ ๐ฒ๐ฐ๐๐๐ถ๐ผ๐ป ๐ฝ๐น๐ฎ๐ป๐, which are different ways of executing the same query. Each execution plan specifies the order in which the tables should be accessed, the join methods, and any filtering or sorting operations. The optimizer then assigns a ๐ฐ๐ผ๐๐ to each execution plan based on the number of disk reads and the CPU time required to execute the query.
Finally, the optimizer ๐ฐ๐ต๐ผ๐ผ๐๐ฒ๐ ๐๐ต๐ฒ ๐ฒ๐ ๐ฒ๐ฐ๐๐๐ถ๐ผ๐ป ๐ฝ๐น๐ฎ๐ป with the lowest cost as the optimal execution plan for the query. This plan is then used to execute the query.
Check in the image the ๐ผ๐ฟ๐ฑ๐ฒ๐ฟ ๐ถ๐ป ๐๐ต๐ถ๐ฐ๐ต ๐ฆ๐ค๐ ๐พ๐๐ฒ๐ฟ๐ถ๐ฒ๐ ๐ฟ๐๐ป.
#technology #softwareengineering #programming #techworldwithmilan #sql
Realized today that `.flat().map()` is not the same as `.flatMap()`, because `flatMap` is more like `.map().flat`, which makes it `.mapFlat()`. Should've gone with `.smoosh()` to avoid confusion.
Did You Know? - In #Kotlin, you can use backticks to include unconventional characters in identifiers. Emojis can make your code even more colorful! ๐
(P.S. - Please don't do this in real projects, of course!)
Eine unendliche Anzahl an Mathematiker trifft sich in der Kneipe. Der erste bestellt 1 Bier. Der zweite 1/2. Der dritte 1/4. Der vierte 1/8. Der Wirt zapft entnervt zwei Bier und sagt:"Macht das mal unter euch aus."