A holistic roadmap for transitioning into Industry 4.0, emphasizing clarity, strategy, and leadership alignment while balancing immediate actions with long-term vision, ensuring organizations navigate this technological evolution efficiently and effectively.
Microblog @antgrasso
Within the constellation of #DLTs, each technology caters to the needs of different industries by offering distinctive features for creating value in a decentralized manner. Delve deeper into a comprehensive understanding of DLTs, read @DeltalogiX report > https://t.co/kSnzv8pGjY
Explaining Sessions, Tokens, JWT, SSO, and OAuth in One Diagram. The method to download the high-resolution PDF is available at the end.
Understanding these backstage maneuvers helps us build secure, seamless experiences.
How do you see the evolution of web session management impacting the future of web applications and user experiences?
Subscribe to our newsletter to download the 𝐡𝐢𝐠𝐡-𝐫𝐞𝐬𝐨𝐥𝐮𝐭𝐢𝐨𝐧 𝐏𝐃𝐅 𝐨𝐟 𝐭𝐡𝐢𝐬 𝐝𝐢𝐚𝐠𝐫𝐚𝐦. After signing up, find the download link on the success page: https://t.co/65PhIVlGbS
Advances in #AI are not just about algorithms; they're about augmenting human potential. Let's embrace the future, intelligently. 🤖 By @ingliguori#DigitalTransformation#Kenovy
What tech stack is commonly used for microservices?
Below you will find a diagram showing the microservice tech stack, both for the development phase and for production.
▶️ 𝐏𝐫𝐞-𝐏𝐫𝐨𝐝𝐮𝐜𝐭𝐢𝐨𝐧
🔹 Define API - This establishes a contract between frontend and backend. We can use Postman or OpenAPI for this.
🔹 Development - Node.js or react is popular for frontend development, and java/python/go for backend development. Also, we need to change the configurations in the API gateway according to API definitions.
🔹 Continuous Integration - JUnit and Jenkins for automated testing. The code is packaged into a Docker image and deployed as microservices.
▶️ 𝐏𝐫𝐨𝐝𝐮𝐜𝐭𝐢𝐨𝐧
🔹 NGinx is a common choice for load balancers. Cloudflare provides CDN (Content Delivery Network).
🔹 API Gateway - We can use spring boot for the gateway, and use Eureka/Zookeeper for service discovery.
🔹 The microservices are deployed on clouds. We have options among AWS, Microsoft Azure, or Google GCP.
🔹 Cache and Full-text Search - Redis is a common choice for caching key-value pairs. ElasticSearch is used for full-text search.
🔹 Communications - For services to talk to each other, we can use messaging infra Kafka or RPC.
🔹 Persistence - We can use MySQL or PostgreSQL for a relational database, and Amazon S3 for object store. We can also use Cassandra for the wide-column store if necessary.
🔹 Management & Monitoring - To manage so many microservices, the common Ops tools include Prometheus, Elastic Stack, and Kubernetes.
Over to you: Did I miss anything? Please comment on what you think is necessary to learn microservices.
–
Subscribe to our weekly newsletter to get a Free System Design PDF (158 pages): https://t.co/FIzCeaWsZV
Current economic and political turbulence could presage the start of a new era structurally very different with a new narrative of progress.
Link >> https://t.co/Lq9Ix7vQL2 @McKinsey via @LindaGrass0#CEO#CIO
How do we transform a system to be Cloud Native?
The diagram below shows the action spectrum and adoption roadmap. You can use it as a blueprint for adopting cloud-native in your organization.
For a company to adopt cloud native architecture, there are 6 aspects in the spectrum:
1. Application definition development
2. Orchestration and management
3. Runtime
4. Provisioning
5. Observability
6. Serverless
Most companies start from Step 1 containerization and gradually adopt CI/CD, service orchestration. This microservice architecture significantly increases the number of instances to manage, so systematic testing and monitoring are required to increase plant observability.
In fact, a lot of companies stop at Step 4 without moving to service mesh and cloud-native networking due to the complexity and the required DevOps talent.
Over to you: Where does your system stand in the adoption roadmap?
Reference: Cloud & DevOps: Continuous Transformation by MIT
Redrawn by ByteByteGo
–
Subscribe to our weekly newsletter to get a Free System Design PDF (158 pages): https://t.co/uc5M7Cdq84
The Ongoing Case For Open Source LLMs
Custom LLMs, long context, and efficient inference
Some folks believe that training open-source LLMs is a losing battle and a complete waste of time.
They argue that the gap between closed models like GPT-4 and open models like Llama will widen and these open-source models may never catch up.
Yes, closed models like Google's Gemini or Open AI's Gobi promise to be way more powerful than GPT-4, so what hope does open source have?
To start with inference on GPT-4 is very expensive. These very large models may be performant but aren't cost-effective. At Abacus, we routinely use fine-tuned versions of LLama-2 and smaller models when we need to run 1M+ API calls a day for standard enterprise applications Q/A, summarization, and NLP at scale. GPT-4 would cost > $100K a day in these cases.
Instruct-tuned LLMs can match the performance of GPT-4 for a specific task. Instruction tuning is a technique that aims to improve the capabilities and controllability of LLMs. It involves further training these models on a dataset consisting of (instruction, output) pairs in a supervised manner. This bridges the gap between the next-word prediction objective of LLMs and the users' objective of having LLMs adhere to human instructions.
For example, we have instruct-tuned open-source models for tasks like Q/A, NER, and classification. These instruct-tuned models are better at generalizing the task to new data and can do so in a resource-efficient manner.
Another shortcoming of currently available closed models like GPT-4 is that they have relatively short context lengths. 8K tokens are default and this means that you can't pass it large documents and ask it to extract the results from there.
Luckily the open-source community has been busy solving practical problems like this. Earlier this week, the paper LongLoRA introduced an ultra-efficient fine-tuning approach to significantly extend the context windows of pre-trained LLMs.
LongLoRA adopts LLaMA2 7B from 4k context to 100k, or LLaMA2 70B to 32k on a single 8x A100 machine, and basic implementation only takes 2 lines of code.
Open-source LLMs have been the focus of the GPU-poor and constraining resources typically have magical effects - efficient, elegant, and simple innovations that solve the problem!
Open-source LLMs have emerged as cheap and efficient alternatives for enterprise AI use cases and will continue to play an important role in the space.
Some have argued that open-sourcing LLMs is dangerous and they may be misused by bad actors.
There is no historical precedent for this argument.
Traditionally, open-source technology has spurred innovation, transparency, and the creation of safe and robust systems. Linux, triumphed over Unix in the OS world, largely because it is open-source and has a huge developer community.
Open source promotes collaboration, community oversight, rapid iteration, and benchmarking all essential for responsible AI development. Open-source developer communities tend to be great at detecting and plugging vulnerabilities.
Disappointingly, big players like OpenAI (despite their name) and Google, haven't open-sourced a lot of their technology. Luckily for the open-source community, Meta has created accessible open-source LLMs. In spite of Meta open-sourcing the powerful 70B LLama-2., the doomsday scenarios outlined by the anti-open-source crowd haven't come true.
Finally, multimodal LLMs (MLLM) are around the corner and if the GPU-rich won't outsource a MLLM, we can always enhance an existing open-source LLM and convert it into a multi-modal model.
In summary, open-source LLMs play a role in the real-world application of AI and are crucial for the democratization of this technology, transparency, and AI alignment
Data classification is an essential strategy that business leaders must use to protect their organizations. In addition, the use of four levels of data classification increases the likelihood of reporting data breaches.
Source @Capterra > https://t.co/IRC81GdR7b via @antgrasso
HTTP status codes you should know
The response codes for HTTP are divided into five categories:
Informational (100-199)
Success (200-299)
Redirection (300-399)
Client Error (400-499)
Server Error (500-599)
These codes are defined in RFC 9110. To save you from reading the entire document (which is about 200 pages), here is a summary of the most common ones.
–
Subscribe to our weekly newsletter to get a Free System Design PDF (158 pages): https://t.co/FIzCeaWsZV
Made a simple visual guide to help everyone understand the key considerations when designing or using caching systems.
- What is a cache
- Why do we need cache
- Where is cache used
- Cache deployment
- Distributed cache
- Cache replacement and invalidation
- Cache strategies
- Caching challenges
- And more.
–
📩 We will write more in-depth articles on these topics. Subscribe to our newsletter so you won't miss out: https://t.co/uc5M7CdXXC
How is an SQL statement executed in the database?
The diagram below shows the process. Note that the architectures for different databases are different, the diagram demonstrates some common designs.
Step 1 - A SQL statement is sent to the database via a transport layer protocol (e.g.TCP).
Step 2 - The SQL statement is sent to the command parser, where it goes through syntactic and semantic analysis, and a query tree is generated afterward.
Step 3 - The query tree is sent to the optimizer. The optimizer creates an execution plan.
Step 4 - The execution plan is sent to the executor. The executor retrieves data from the execution.
Step 5 - Access methods provide the data fetching logic required for execution, retrieving data from the storage engine.
Step 6 - Access methods decide whether the SQL statement is read-only. If the query is read-only (SELECT statement), it is passed to the buffer manager for further processing. The buffer manager looks for the data in the cache or data files.
Step 7 - If the statement is an UPDATE or INSERT, it is passed to the transaction manager for further processing.
Step 8 - During a transaction, the data is in lock mode. This is guaranteed by the lock manager. It also ensures the transaction’s ACID properties.
–
Subscribe to our weekly newsletter to get a Free System Design PDF (158 pages): https://t.co/FIzCeaWsZV
Why is Kafka fast?
There are many design decisions that contributed to Kafka’s performance. In this post, we’ll focus on two. We think these two carried the most weight.
1️. The first one is Kafka’s reliance on Sequential I/O.
2️. The second design choice that gives Kafka its performance advantage is its focus on efficiency: zero copy principle.
The diagram below illustrates how the data is transmitted between producer and consumer, and what zero-copy means.
🔹Step 1.1 - 1.3: Producer writes data to the disk
🔹Step 2: Consumer reads data without zero-copy
2.1: The data is loaded from disk to OS cache
2.2 The data is copied from OS cache to Kafka application
2.3 Kafka application copies the data into the socket buffer
2.4 The data is copied from socket buffer to network card
2.5 The network card sends data out to the consumer
🔹Step 3: Consumer reads data with zero-copy
3.1: The data is loaded from disk to OS cache
3.2 OS cache directly copies the data to the network card via sendfile() command
3.3 The network card sends data out to the consumer
Zero copy is a shortcut to save the multiple data copies between application context and kernel context.
–
Subscribe to our weekly newsletter to get a Free System Design PDF (158 pages): https://t.co/FIzCeaWsZV
This is the flowchart of how slack decides to send a notification.
It is a great example of why a simple feature may take much longer to develop than many people think.
When we have a great design, users may not notice the complexity because it feels like the feature just working as intended.
What’s your takeaway from this diagram?
Image source: Slack eng blog
–
Subscribe to our weekly newsletter to get a Free System Design PDF (158 pages): https://t.co/uc5M7CdXXC
Generative AI inspires and worries the working world. However, PwC found that more than half of workers see benefits in productivity and opportunities, while only a third express concerns.
Source @PwC Link https://t.co/7wI2MHeyd4 via @antgrasso#GenerativeAI#AI
From #ExploreSAS - Generative AI is reshaping industries, with McKinsey highlighting its significant revenue potential. SAS Viya is at the forefront, blending innovation with real-world application.
More > https://t.co/XymEYhRpSe
Paid partnership w/ @SASsoftware#SASVisionary