The #WorldSeries finale is simply the start of a new season! At @SchoolhouseData, we're excited to continue story-telling, and start building our own forecasts. We'd be thrilled to hear from you.
#MLB#startups
@kelseyhightower Reminds me also of Nick Schrock's Dagster blog: "LLMs and AIs cannot author software unsupervised because human language is not precise enough to specify requirements."
You know that laugh or "*yes*" that comes out when you hear from someone who just "gets it"?
Thank you @jthandy for your sharp insights.
- https://t.co/0A6xupAQEq
𝗪𝗵𝗮𝘁 𝗶𝘀 𝘁𝗵𝗲 𝗺𝗼𝘀𝘁 𝗶𝗺𝗽𝗼𝗿𝘁𝗮𝗻𝘁 𝘁𝗵𝗶𝗻𝗴 𝗶𝗻 𝗰𝗼𝗺𝗽𝘂𝘁𝗲𝗿 𝘀𝗰𝗶𝗲𝗻𝗰𝗲?
There are different thoughts on this. Is it problem-solving? Or maybe data structures and algorithms? Some think it is architecture.
According to Professor Of Computer Science John Ousterhout from Stanford University, it is a 𝗽𝗿𝗼𝗯𝗹𝗲𝗺 𝗱𝗲𝗰𝗼𝗺𝗽𝗼𝘀𝗶𝘁𝗶𝗼𝗻. In his book "A Philosophy of Software Design," he advocates that good software design means fighting complexity and recommends different strategies for fighting complexity. Here are some which I agree with:
𝟭. 𝗖𝗹𝗮𝘀𝘀𝗲𝘀 𝘀𝗵𝗼𝘂𝗹𝗱 𝗯𝗲 𝗱𝗲𝗲𝗽 𝗮𝗻𝗱 𝗶𝗻𝘁𝗲𝗿𝗳𝗮𝗰𝗲𝘀 𝘀𝗶𝗺𝗽𝗹𝗲
The famous David Parnas paper in 1970 ("On the Criteria To Be Used In Decomposing Systems into Modules") mentions that we need simple interfaces but a lot of functionality inside a method or class. We could typically see that we have shallow methods with only one line of code inside. He considers the biggest mistake too small and too shallow of classes. As an example, we can see (something he calls classisitis), in Java we need two to three classes to read a simple file: 𝙵𝚒𝚕𝚎𝙸𝚗𝚙𝚞𝚝𝚂𝚝𝚛𝚎𝚊𝚖→𝙱𝚞𝚏𝚏𝚎𝚛𝚎𝚍𝙸𝚗𝚙𝚞𝚝𝚂𝚝𝚛𝚎𝚊𝚖. Bigger modules and more generic interfaces and classes encourage information hiding and reduce complexity.
In addition, the author mentions a famous API design antipattern involving overexposing internals, which later adds to an architecture debt.
𝟮. 𝗠𝗶𝗻𝗱𝘀𝗲𝘁 (𝗦𝘁𝗿𝗮𝘁𝗲𝗴𝗶𝗰 𝘃𝘀 𝘁𝗮𝗰𝘁𝗶𝗰𝗮𝗹 𝗽𝗿𝗼𝗴𝗿𝗮𝗺𝗺𝗶𝗻𝗴)
Most people use a tactical approach, where the goal is to make something work. However, the result is a bad design with a lot of tech complexity, which usually results in spaghetti code. Complexity is not a single line but many lines in a project, which we overlook as a whole. We sometimes also create a feature that we may need (YAGNI). His recommendation is to take a strategic approach where the working code is not the only goal, but the goal should be great design, which simplifies development and minimizes complexity.
𝟯. 𝗜𝗻𝘃𝗲𝘀𝘁 𝘆𝗼𝘂𝗿 𝘁𝗶𝗺𝗲 𝗶𝗻 𝗾𝘂𝗮𝗹𝗶𝘁𝘆
What we usually see with startups is pressure to build products quickly, which can result in spaghetti code. But good code gives you more advantages. One of those is to hire great people who like to work with high-quality codebases. In addition, we need to make our development 10-20% slower, and you will get it all back in the future (this reduces tech debt also). Code and refactor in small steps.
I can't entirely agree with the usage of 𝗰𝗼𝗱𝗲 𝗰𝗼𝗺𝗺𝗲𝗻𝘁𝘀, although the author mentions that we need to use them only for what is not apparent. He also disagrees with the 𝘂𝘀𝗮𝗴𝗲 𝗼𝗳 𝗲𝘅𝗰𝗲𝗽𝘁𝗶𝗼𝗻𝘀, where the point is that they are a massive source of complexity. Exceptions are usually good, especially if they are appropriately monitored and logged.
Do you agree?
#softwaredesign #softwarearchitecture #programming
Grateful for the #python `pandas` books trifecta:
- effective pandas by @__mharrison__
- python data science handbook by @jakevdp
- python for data analysis by @wesmckinn
Thank you for all the hard work ... your impact and knowledge are inspiring!
Grateful for the #Python pandas fundamentals in Effective Pandas 2 by @__mharrison__ ! Helpful focus on core data types and the numpy -> pyarrow types backend transition. All context not directly covered in Statistics curricula.
It's now much faster to start a new #datascience project, thanks to the "cookiecutter" tool (https://t.co/wti3V5g8aJ). Standard structure, thoughtfully designed.
Thank you @drivendataorg!! Data science standards can be elusive, major pain point.
"LLMs and AIs cannot author software unsupervised because human language is not precise enough to specify requirements."
- @schrockn (https://t.co/VSfQ6Hb401)
Great coverage of the coder-but-not-engineer landscape.
Who's a better addition to your portfolio: Brewers' Joey Ortiz or Cardinals' Masyn Winn? We think Ortiz has a better risk-reward profile (chart).
These guys have similar average performance ("expected return"), but more volatility in Winn's daily stats.
@CardPurchaser
Stocks and trading cards are similar: all about risk-reward. We built risk-reward profiles for 2024 #MLB hitters, from daily performance data (chart).
The labeled players:
1. Heliot Ramos
2. Gunnar Henderson
3. Aaron Judge
4. Colt Keith
5. Javier Baez
@CardPurchaser
@Buddahtime44 @SchoolhouseData @CardPurchaser right! Suppose that every Friday, updated 2024 WAR forecasts landed in your inbox. How would you like to use those?
𝗕𝗼𝗼𝗸𝘀 𝗘𝘃𝗲𝗿𝘆 𝗦𝗼𝗳𝘁𝘄𝗮𝗿𝗲 𝗘𝗻𝗴𝗶𝗻𝗲𝗲𝗿 𝗠𝘂𝘀𝘁 𝗥𝗲𝗮𝗱 𝗶𝗻 𝟮𝟬𝟮𝟰.
You probably already noticed that I'm a big fan of reading. I usually read 3-4 books per month. There are two ways to learn from knowledgeable people: to work directly with them or to read what they have written. The first option is the best, yet it is often impossible to do. So, we have books written by people who are probably the best at this in the world at the time of writing.
If we look at the software engineering world, there are many gems here, but I will recommend the best books per area of work. These books will help you not only to become good at specific technology but to become a great software engineer overall.
𝟭. 𝗚𝗲𝗻𝗲𝗿𝗮𝗹:
🔹 The Pragmatic Programmer by David Thomas and Andrew Hunt (https://t.co/NCSpr3ZhGb)
🔹 Code Complete: A Practical Handbook of Software Construction (https://t.co/OXMYyabHma)
🔹 Modern Software Engineering by David Farley (https://t.co/2X5eWCHcni)
🔹 Software Engineering at Google (Free - https://t.co/GXZxoCbrva)
𝟮. 𝗖𝗼𝗱𝗶𝗻𝗴 𝗽𝗿𝗮𝗰𝘁𝗶𝗰𝗲𝘀:
🔹 Clean Code by Uncle Bob Martin (https://t.co/4Ml52XBKKb)
🔹 Head First Design Patterns by Eric Freeman (https://t.co/4jXkPd8vcK)
🔹 Refactoring by Martin Fowler (https://t.co/8fbR93LNy0)
𝟯. 𝗗𝗮𝘁𝗮 𝘀𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗲𝘀 𝗮𝗻𝗱 𝗮𝗹𝗴𝗼𝗿𝗶𝘁𝗵𝗺𝘀:
🔹 Grokking Algorithms by Aditya Bhargava (https://t.co/1q2hWfaONO)
𝟰. 𝗗𝗮𝘁𝗮:
🔹 Learning SQL by Alan Beaulieu (Free - https://t.co/44iy3v91JQ)
𝟱. 𝗧𝗲𝘀𝘁𝗶𝗻𝗴:
🔹 Growing OO Software by Tests by Steve Freeman (https://t.co/joi6Q8nm4W)
🔹 TDD by Example by Kent Beck (https://t.co/IxVGfJymQu)
🔹 Unit Testing Principles, Practices, and Patterns by Vladimir Khorikov (https://t.co/7VyFPkUpZS)
🔹 The Art of Unit Testing by Roy Osherove (https://t.co/rqNoqJH49t)
𝟲. 𝗔𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲:
🔹 Fundamentals Of Software Architecture by Mark Richards and Neil Ford (https://t.co/LOnF7783bl)
🔹 A Philosophy of Software Design by John Ousterhout (https://t.co/rBeKHE0w6P)
🔹 Clean Architecture by Uncle Bob Martin (https://t.co/fojqHumHo3)
🔹 Domain-Driven Design Distilled by Vaughn Vernon (https://t.co/pT8GZQmrR5)
🔹 Software Architecture the Hard Parts (https://t.co/K7AqvDOoSN)
𝟳. 𝗗𝗶𝘀𝘁𝗿𝗶𝗯𝘂𝘁𝗲𝗱 𝘀𝘆𝘀𝘁𝗲𝗺𝘀:
🔹 Understanding Distributed Systems by Roberto Vitillo (https://t.co/NmApvbFyQj)
🔹 Designing Data-Intensive Applications by Martin Kleppman (https://t.co/2Rtjfs987o)
𝟴. 𝗗𝗲𝘃𝗢𝗽𝘀:
🔹 DevOps Handbook by Gene Kim (https://t.co/EHmuVxZrKi)
🔹 Continuous Delivery by Jez Humble and David Farley (https://t.co/tePOX3hfs3)
🔹 Accelerate by Nicole Forsgren (https://t.co/HqHfEjAmB4)
𝟵. 𝗧𝗲𝗰𝗵-𝘀𝗽𝗲𝗰𝗶𝗳𝗶𝗰:
🔹 C# in Depth by Jon Skeet (https://t.co/u4M31XN673)
🔹 Effective Java by Joshua Bloch (https://t.co/5yc8S6sMS3)
🔹 Fluent Python (https://t.co/IO5IFIuXOb)
𝟭𝟬. 𝗠𝗮𝗰𝗵𝗶𝗻𝗲 𝗹𝗲𝗮𝗿𝗻𝗶𝗻𝗴:
🔹 The Hundred-Page Machine Learning Book (https://t.co/31A2XuGKSR)
🔹 Designing Machine Learning Systems (https://t.co/8hXFovtTzU)
𝟭𝟭. 𝗟𝗲𝗮𝗱𝗲𝗿𝘀𝗵𝗶𝗽:
🔹 The Five Dysfunctions of a Team by Patrick Lencioni (https://t.co/bKS3xhjCQv)
🔹 Drive by Daniel Pink (https://t.co/GVyKMoAkUu)
🔹 The Making of a Manager by Julie Zhuo (https://t.co/BiiOaJFmKz)
𝟭𝟮. 𝗣𝗲𝗿𝘀𝗼𝗻𝗮𝗹 𝗱𝗲𝘃𝗲𝗹𝗼𝗽𝗺𝗲𝗻𝘁:
🔹 How to Win Friends & Influence People (https://t.co/zKnkSvI4hv)
🔹 Deep Work (https://t.co/bOLnphk2Z4)
#softwareengineering
@statwonk haha, well played! Still, I think there's good truth in your quip. Like you've said -- if use case is insensitive to model complexity, that's critical.
@statwonk have recently thought - maybe I struggle to justify complex models because, I lack vision to connect to use cases. Probably part of the explanation. But maybe the other part is structural, like you've put it.