Learn how to build and train a GLM 5.3 flash model with 25 million parameters from scratch using just a regular CPU. This tutorial covers basic pre- training and reinforcement learning techniques, teaching you to design effective experiments and think like a modern AI researcher. Let's build and train GLM 5.3 flash, also known as ox alpha, the newest and most popular model from Google AI. We will perform both pre- and post-training, as well as build everything from scratch. So, you'll learn to do all the things that GLM, OpenAI, and Anthropic researchers do, just on a simpler scale. The work of an AI researcher has changed dramatically in just a few months. So this is the most relevant course you will need if you want to become an AI researcher. All code is available on GitHub. These are relevant experiments. We will reproduce everything you see on the charts. Link is under the video. This is the latest job posting from Anthropic, and it describes 95% of AI researcher roles. First, in this course, you will understand architectures and algorithms. Then I'll show you how to do real AI research: design environments, reinforce learning, and ask questions. Now Claude, Code Codex, and other AIs write the code for all experiments and can do so autonomously. Your job as an AI researcher is just to have a general idea of what experiments to run. Even if Anthropic and OpenAI could invent new algorithms, right now it's transformers and reinforcement learning. I think they focus 99% or more of their efforts on developing data, environments, and then optimizing these algorithms, which you'll understand in this video. Thinking about the future, the work of an AI researcher will be mostly about generating high-level ideas, not programming. AI models will do all the programming and experimentation. But you still need to understand how these algorithms work. You don't have to write code by hand, but it is people who can invent better, faster concepts, faster algorithms, and create more sophisticated system designs. And then AI will be able to implement, test, and test everything. That's why we'll look at the code and understand how it works, even if we don't have to write everything from scratch manually. But if you want to invent new algorithms as an AI researcher, you need to understand how modern algorithms work. So while the original model is incredible and has billions of parameters, our model will only have 25 million and can be trained on a CPU, a graphics card, a MacBook, or anything else. I made it very simple. And that's what you want to do. You can still do so much research here. You can learn to ask questions, identify goals, topics, and research directions for AI. Even if you have a lot of computing power, you won't be able to do more with it than just on the processor. I don't think it's necessary unless you want to create a model that others will use. But here we are trying to learn. So, we want to learn how to ask questions, and AI can very easily help answer them. AI will become increasingly capable of answering questions, and for now, humans will still be asking those questions. Of course, we don't know what will happen in a year: superintelligence or something else, then humans will become unnecessary , and AI will do everything itself. So for now, we're just focusing on how to ask questions, research questions , define goals and directions. Here are some examples of the questions we will ask and answer in this study. So these are real things you can do. You can learn to conduct research on just the processor. Let's say you are training reinforcement learning for reasoning. Should we reward an imperfect score, a partial score, or not? Well, we'll try it and see what works better. How severe should the penalty be for invalid code? And also the temperature or randomness of the generation of the next word in reinforcement learning. It's a balance between research and using what is already known. These are just a few examples, and it is your job to ask such questions. This is what distinguishes an experienced researcher from a novice. And AI can now perform all these experiments on its own. Even Soul Opus and Astra Fable, anyone can conduct such experiments. By the way , if you need help asking these questions and want to join a real AI lab, join our community where we simulate the work of a real AI research lab. You will either set your own research goals or use ours , answering our questions and helping us find new ones. We simulate the learning process of humans, training them how to conduct real AI research. Okay , so let's get started. We know that LLMs predict the next word, token, or symbol (in our case) because we create, and you will create, a small model. As I said , we only need small models with fast iteration. So in this example, the input could be return X plus and the model would predict one depending on the task or query. This is the next token. True large language models use BPE tokenizers, tokenizers based on byte pairwise encoding. So, if you have text like that, they will break it down into subwords or individual tokens. They usually have 150,000 or 300,000 such subwords, and can encode any text. Any text can be expressed through this dictionary. But to make the process clearer for learning, we will take each ASCII character. This way, each letter will become its own token. If we have a def function, int or something else, each character will be a separate token. So, our dictionary will only have 260 tokens. Therefore, we don't need to teach tokenization. It's easier. Plus, it solves a big problem that would have arisen later. Some of you are already familiar with the architecture of transformers and LLM, where there is an embedding layer and an output layer. Both of these weight matrices use the token dictionary as one of their dimensions . If we have huge dictionaries of 250,000 tokens, then in a very small model 95% of the parameters will simply go to dictionary transformations. And only 5% of the parameters will remain for the layers themselves and the model as a whole. If you didn't understand what I said, that's okay. It's just that small models require a small dictionary. A dictionary of 260 tokens is good. So, here's how the next token is generated in the transformer. Let's say we have some input. It could be a number, a query, or anything. And each of these tokens in the input has a corresponding vector embedding that represents it. Each token is replaced by its own vector embedding, which is a sequence of numbers encoding the value of that token. Token A has its own numbers, token B has its own, plus has its own, minus has its own. The vector for "plus" seems to explain or contain information about how addition is performed and what the functionality of this token is. So each token in the query is replaced by a corresponding vector describing its meaning. These vectors are learned by the language model during its training. So, LLM learns to represent the concept of "plus" using some numbers. Then we have a sequence of vectors—this is our prompt, which goes into the transformer, into the layers. It is processed by a transformer. And at the end, we want to predict the next token. So , it simply transforms everything that was processed, the entire prompt, into a probability distribution over the entire dictionary. That is, based on the prompt, you will get for each token from the dictionary the probability that it will be next. That's roughly how it works. This is the basic idea of a transformer. Now, GLM 5.3 flash has new architectures, new components that we haven't talked about on my channel before. So, a hybrid of linear and sparse attention. Then a mix of experts—we've already discussed this, it's a very well-known thing. Next — deep six, manifold-bounded residual relations, hyperlinks. And another quite well-known and old thing: using a single matrix to convert the input and output dictionaries into embeddings. So, entry and exit. These are the specifications of our model. The original is huge, and ours is tiny. It's layers, many layers, but far fewer experts. The context is only 192 tokens, while GLM has 1 million tokens. And you will see that in our experiments before reinforcement learning, the model will not perform any task successfully. And after reinforcement learning, she will successfully complete 16 out of 24 tasks. These will be very simple tasks, such as writing a Python function that multiplies an input value by two. But you will see a difference before and after reinforcement learning. You will learn how to ask questions, develop ideas, and environments for learning the model. You will also be able to see how the model learns during training. This is a graph that shows the learning process during training. Now let's look at the code. We're going to understand algorithms, how they work, but no one is going to write it from scratch. But it is people who should understand the algorithms and then improve them. Tell the AI what to try, and then the AI will try out many ideas based on yours. But you have to know the direction. So, that's what we're going to analyze in this part of the code now. First you have the embedding and then the login ID is passed. Let's say this is our token ID. This will simply look in the dictionary and select the vector I mentioned earlier. So this can be a huge matrix if you have a huge vocabulary. And this can be a huge, powerful matrix. And if you have a small model, then this will be 95% of the parameters. That's why for small models we use a small dictionary, bytes. So, next it will simply extract the vector of this token. And this happens at the very beginning, when the request is processed. So we have, after we convert all these vectors, we will have the dimension of the vector, as well as the number of tokens. So the number of tokens, the dimension of a token, and the batch is the number of independent conversations in this huge X- vector, the X-tensor. You see how it just takes this X-vector and creates an array of these X- vectors. These will be residual connections. Subsequently, each of the residual links is a hyperlink with a restriction on the manifold from DeepSeek and ByteDance. And each of them will pass different information on without processing it through the attention mechanism, so it won't get mixed up, it'll just be passed on like highways where you pass different types of information. I guess I wo n't go into too much detail here. If you're interested, I made some videos and you can read the article because I think it's getting boring and I can't help you much with that. You need to practice it yourself, read it, understand it, and try it. If I just explain, you might not be able to follow the thought. So let's just go through these things. These hyperlinks will be mixed together depending on certain weights and mixing methods. So they exchange information with each other, between attention mechanisms, feed- forward neural networks, and everything else. And in the end, they merge into a single vector of the hidden dimension. So, thanks to these hyperlinks, you don't just have one residual connection. You allow the transformer to shape and pass through many different types of data during processing. And then this hidden state is converted into a probability distribution over the entire dictionary using a certain weight matrix. That is, the distribution of which token is most likely to be next and best fits this context. And this is another part that can be extremely huge compared to the rest of the model if the dictionary is large and the model is small. So this weight matrix that turns the hidden state into a dictionary can be quite large. And if you remember, at the beginning of Transformer we also had a hidden state dictionary, where the hidden state is actually a vector embedding. So you can actually use this same weight matrix for both cases. So, you can use the same matrix to choose a vector for each token and create a probability distribution at the end, and depending on the architecture, this works well. This saves a lot of parameters that can be used elsewhere. Another part is RMS normalization. It simply normalizes the vectors. You see how the numbers here are very different from the others. In neural networks, due to the way error backpropagation and mathematical calculations work, this can interfere with calculations , cause weights to explode, making them insignificant, or one value can drag all the weight onto itself, making the others unimportant due to excessive magnitude. So you usually want the numbers to be in the range of -1 to 1, close to zero, and similar in scale. Then these numbers won't " go crazy" when you multiply them many times in a row. This is positional encoding. How does the Transformer know which token comes after another? Because in the attention mechanism, tokens are multiplied by other tokens, you do not know the order of the tokens in the query. Therefore, each token vector is returned. For example, if you have a vector, you take pairs of dimensions and return each of them. Depending on where this pair is in the sequence and where the token stands in the token chain, you get slightly different twists. So, the Transformer will learn the position of the vectors based on how much they are rotated. In GLM they use NOPE. This is a version of RoPE in the main model, as well as RoPE in the sparse indexer. This sparse indexer checks in the attention mechanism which part is worth paying attention to and which is not, because here attention does not work in such a way as to cover everything in a row or the entire context window. So this token will only pay attention to the most relevant tokens. This indexer will look for the most relevant tokens so that the current token can pay attention to them. An indexer is also an attention mechanism, but much easier, cheaper, and faster. And it simply looks for relevance and doesn't require as much computation as regular attention. So, it selects the most relevant tokens, and only those are computed using full conventional attention. So for each token in the attention mechanism, it will select the n most relevant tokens. It could be the 2000 most relevant tokens or 1000, and then it applies attention to them. And here, attention scales based only on this small indexer. So this small indexer, instead of scaling like this , will scale much slower because the attention size is always fixed, it doesn't scale, and the indexer scales depending on the length of the sequence. Sequence length. So I'm talking about how computation scales with sequence length. And it will scale much more slowly with sequence length, requiring significantly less computation, because this sparse indexer is a much lighter type of attention and does fewer computations. He only needs to find the relevant tokens. As you can see, linear attention is different from this standard quadratic attention. It keeps, so as not to confuse you, let me first explain that their attention mechanism does not contain positional encodings. And only a sparse indexer contains positional encodings. So, "nope" means no positional encodings. We use "rope" in the attention and indexer to simplify our model. And linear attention is just linear attention KDA. So , this is a state space model . I'll explain it. So, we have linear, linear, linear, and sparse layers. Linear layers are cheap. They always maintain a constant vector. Imagine you have tokens, there is a state like a matrix that serves as a memory for all the tokens. And all the tokens that come in are fed into this state , into this memory. And it is always the same size. It's like just adding to it. So the linear layers will remember recent tokens. They will not remember past tokens. Because the further back a token is, the more it has been changed, the more it has been distorted by newer tokens. So, that's it for recent tokens. And they are so fast, much faster than attention. And attention is intended for long-lived tokens anywhere in the context window. In my experience, this trains much faster and performs inference faster. Unlike attention. That's why it's optimized for inference. So, linear layers cheaply carry compressed history, memory. Sparse layers periodically restore precise, distant details. Here is the code for the layers and the attention mechanism. So let me show you, it's very simple. You simply take every fourth layer and set it as sparse focus. And this is linear attention. So you have some tokens, and they all go into this memory matrix. This is called the current key- value state. So as the context grows, the length of the sequence increases, this state remains unchanged. This is a way to store data without quadratic scaling. The trade-off is that you lose some information due to compression. So , with sparse attention, you typically look at a fixed window of the last n tokens, for example, the last 2000 tokens. And then you can browse at regular intervals or search for relevant tokens behind. Next we have a mix of experts (MoE). For example, in GLM you have 288 selected experts. And for each token, eight out of 288 will be selected. There is also one shared expert. Each token goes through a common expert who studies common information for all tokens so that other experts do not waste their resources on it . So in our small model, we can have eight routed experts, from which we choose two, plus one shared. I can also have a model with only two experts and one routed. That is, the chosen one. Here I show the configuration of our model for training. We've already mentioned this: eight experts, 12 layers, etc. So there are four hyperlinks. Next, you have them within attention, that is, between attention and MoE . They seem to merge into attention, then expand, merge again into MoE, and so on. Here they are connected as residuals. And here they are also connected as residues, so they pass through the MoE. So, attention moves information between token positions. MoE applies specialized transformations to each position. Residual attention preserves the previous representation by adding both updates, and multiple connections preserve different types of information. The input and output share one common matrix, so the output weights and the embedding weights are the same matrix, as I mentioned earlier. That is, reading byte tokens and predicting the next byte use the same weight matrix. One matrix studies both token embeddings and output logits. Or better said, a matrix of weights that converts them into these values. So weight tying , as it's called, eliminates the second large matrix and forces reading and predicting bytes to use a single learned geometry. So, maybe there is some benefit to this too. I don't know if there is research, I guess there is, maybe you can do something too. So, GLM flash also has vision, that is, there is an image, it transforms this image into certain patch embeddings, converts them into tokens , and then mixes them with the main model. For images, it's divided into patches, each patch is a token, so it has these 2D positional rotational embeddings. Similar to rope, just 2D. Let's say you have image number eight, it's a 32 by 32 RGB image. And in this case it's going to be divided into 8 by 8 patches. So in our small version that we coded, there are two vision blocks, by the way, all of this code is on GitHub. All the code is right here on GitHub, so you can check it out. Here are these experiments. So each of these patches, four patches, will become one vector, one token embedding vector. That is, merging 2 by 2. And so , as a result, we will get 4 by 4, that is, 16 tokens from this 32 by 32 image. So, if you have another image, not 32 by 32, you need to resize it to this size and add or crop it. So this is just for our encoder. GPT has any image encoder, you know, these other models, and we can create any as well. I just made this one really quickly. So, the image tokens will have certain labels, like " image start", then all 16 image tokens, and then " image end". This is how they are added to the context window. There are these marks. And there may be text around it. By the way, I didn't mention that this beginning of the sequence is a special token. In the dictionary, it simply shows the beginning of a conversation sequence. And then you can add text to all of these tokens. And then, next would be, for example, the number seven —the intended token. If this is a picture of the number seven. We taught a small transformer to look at a picture of a number and tell us what number it is. So if you are interested in the details, just ask your Codex or Claude Code to clone this repository. You can ask him about the code, everything, how everything works. You can also have her draw these graphs, like I did here. So, these are examples . Now, of course, we can see it clearly. It's very easy for us to see this. Anyway, our visual encoder, our visual transformer, successfully recognized the numbers using this method that I explained. If you want to do some research , you can make this task much more difficult and then try to see at what point it stops recognizing them. So let's take a look at the previous training. So, we have a certain prompt, which is a function in Python. And then we have this addition that is generated by the model. Uh, because the task is to return x + 1. Now , uh, during the pre- training, he doesn't " think" about this answer. It simply repeats what is in the data. So, he just repeats uh this . That's how he learns. He already has all of this. So it's like simulation learning. He imitates. Now, during pre- training, he simulates the writing process. He studies what all these tokens, all these symbols, mean so he can generate them. And then, during reinforcement learning , he will learn to generate , to think about, what answers are actually correct. So, he learns during post-training, he learns to be smarter. So you have "red", it assumes "you", if you have "return", it assumes a space. You get so many learning examples from just one prompt . This is a classic pre- learning pipeline: probabilities, cross- entropy, assigned weights, and then the optimizer updates all parameters based on the correctness result of the next token. So, AdamW—most people use Neon, but I did it for simplicity, so we don't use both, we just use AdamW. You see the gradients go from an incorrect probability distribution to a more correct one. They literally change the choices. It's going to change millions of weights, so the correct byte becomes a little more likely each time. You see, at the beginning, before the pre- training, it just generated some random characters, if this is it— prompt. But later he learned to generate it. Now let's look at some of the research questions and experiments I conducted. So, this is how research is conducted. This is a very standard example, so you can see that if you have less data and you repeat the data, it will learn faster at first , but later on the larger amount of data will just take over. So, variety is not useful if you only take a few steps, but then, at 200 steps, it becomes more and more useful. So, let's say you have three different types of data: type A, type B, and type C. Which is better: training on type A first, then B, and then C , or alternating them? You will see that if you train the model in this way, then when it starts learning C, it will start forgetting A and B. Therefore, it is better to alternate them, and this can also be proven through experiment. You can do this experiment if you want. Therefore, alternating data will improve generation and also prevent catastrophic forgetting of previously learned information. Homogeneous blocks are likely to amplify bias in favor of recency or forgetting, but you need to know what exactly your experiment is measuring . Here we only measure performance. We don't measure whether the bias in favor of recent data or forgetting is increasing, so I can't make such statements. I can only assume that it is possible. How about this option? What if we first train the model on repetitive patterns, simple patterns, and then train it on more diverse data? Maybe it's worth learning simpler things at first, and then more complex ones. We compare this to only simple and only complex tasks . We already know that simple things show worse results, and training on only the same data leads to overtraining. And here the curriculum has not made significant progress compared to training on complex data from the very beginning. Perhaps this complex data was also too simple, so I need to think carefully about my experiment setup. Also, this difference may not be statistically significant depending on how many times I took the measurements and how . I need to think about all of these things, and AI can help you understand and think about all of these aspects of planning your experiment. Now let's look at reinforcement learning. Reinforcement learning is different from prior learning. Pre- learning is simply learning to repeat text and understand the meaning of each token. In reinforcement learning, the model tries to generate responses , and we reward the correct ones. So somehow—I don't know how—the model learns to think. And if you can figure out how, you'll become the most famous AI researcher. But in reinforcement learning, the model learns to use its “thinking” through the transformer to find the correct answer. Reinforcement learning now tracks a separate deployment process. Completing a single task takes, perhaps even days, at cutting-edge labs like OpenAI and Anthropic. If you want to train a model to work for days, then during training you will let it attempt tasks that last for days. So, by the way, if you want to help these labs, you can also work on the findings. Faster inference means you generate data faster, so you can train the model faster. So, generate code, run tests, evaluate what's right and what's wrong. Now there are different ways to evaluate: you can evaluate only correct or partial results, for example, and then update the model weights. You should not show model tests. You will only evaluate her result because if you show her the test, she may "break" the reward, i.e., for example, just read the result instead of learning how to solve it. Learning to hack Hugging Face or other systems—that's what happened. And during our post-training training on a small example, we went from 4% correct answers to 45% after training. So, usually you can just tell Fable, Opus, and Soul to do this post-workout training for you. This is what you need to learn. What Opus and Fable don't know is how to get the direction of the experiment and the ideas for the experiments. They can perform reinforcement learning simply because they've done it many times online, they know how to do it. But your judgment, right now your judgment about experiments—that's what you need to practice. What ideas? And then she will be able to conduct experiments. So, before reinforcement learning, there were zero out of 24 questions, after reinforcement learning, there were 16 out of 24 questions. "Pass at one" means you only give it one try and then evaluate whether it is correct or not during testing. If you have a "pass at eight" it means you give eight attempts, and if at least one of the eight is correct, you consider it the correct answer. That is, you give the same question and let them try to answer it several times. But here we just let it answer once and then see if it's correct or not. Of course, you always need to look at how exactly you measure your results. For example, in this case, we test her on similar tasks that she learned during training. They are not the same, but they are similar because that way she will learn the patterns, learn to do these tasks, and then just learn to apply them to other numbers or maybe other operations. So, the more variety and versatility you want, the more you need to scale the model. More parameters, more calculations, more data. And then it will become more universal, more diverse. But don't worry about it, you can always do research. I believe that good research can be done even on a processor. It seems to me that if you want to get a job at OpenAI or Anthropic, you don't need GPUs. You want to learn how to create environments for reinforced learning , understand how to think about them . Verifying whether the model learns in these environments is important. It is not necessary to scale unless you are working directly on scaling or pre-training . If the task is to return two multiplied by X, then before reinforcement learning she returned X multiplied by X, and after—X multiplied by two. So, here she generated correctly, she learned. So reinforcement learning increased the probability that the correct response would be generated based on all parameters. There are many ways to structure rewards. For example, you can give a plus one if all tests are passed, a zero for the correct format but the wrong number. And maybe minus one for incorrect Python syntax. This is just one way . You can simply assign a zero for an incorrect or erroneous result , and a one for all tests passed. You can also give a unit if it is written in the correct format. If she writes correct Python code, you can reward her. So, there are many different ways. These are experiments, empirical evidence, but I also think that the data is much more important than the structure of these algorithms. You'll find a good enough algorithm, and then you'll have less of a return on thinking about the algorithm itself compared to thinking about the data. You will greatly improve your model if you think about what exactly you are teaching it. What kind of environment is this? What is she trying to do? Can she do it? Can she learn this? Therefore, tailoring the questions to the environment, to the data, and to what it is she is studying is more important. This is most important when you already have a good enough algorithm. This is why DeepSeek no longer puts significant effort into GRPO. GRPO, a reinforcement learning algorithm , is good enough. If they put in more effort, it won't yield the same return as the investment in data, training, post- training, and environments. The verifier will run hidden tests. You don't show these model tests. Now you have things like GPT, Fable, Astra. They're going to They're probably going to hack and see your, uh, verifiers. So, this is what is happening right now. So, there are different ways you can do this. For example , in GRPO and other algorithms, you can ask the same question and let the model generate 16 answers. And then you can reward each answer depending on whether it is correct, incorrect, contains Python errors, etc. And we want to calculate how much each answer is better or worse than the other answers, the average. That's why they call it an advantage. So, the reward is simply your 1, 0, or -0.1. This is your reward. Minus the average value of other rewards. So, it will simply calculate whether it is better than average, and how much better or worse it is than average. So, better answers will become more likely, and worse answers will become less likely. This is good because you don't need a separate critic, like a separate model, to review the results and evaluate them. You don't need people to do this. You can simply let the model generate multiple times and compare the results with each other. So, I won't update the entire model with these reinforcement learning signals. I will only update the last block, the final normalization layer, and the original "head". These are the associated weights between the input and output, as I mentioned earlier. But only the end, because in this way, firstly, we do not change the previous layers that have other knowledge, and then they may face forgetting other knowledge, because now we are changing them only on this new data. We don't train it on old data. So, this could change things too much . And it's also faster, updating fewer weights. So, you can also use LoRA. You can use different methods. You can update the entire model. So, there are different methods here. You can research this if you want. You can also conduct research and experiments, try out different ways to update multiple blocks and try to understand why. So, try to use toy examples when conducting experiments. And that's what MIT is talking about. MIT says you need to start with small toy examples so you can understand what's going on. This is better than doing huge, complicated experiments where you don't know what's going on. RLOO is the algorithm I use. So, algorithms, as I said, don't really matter. Data is what matters. What he learns depends on what data we give him. Because our model was small, we can only improve it on tasks similar to the ones we train it on. On the other hand, if you have large models like OpenAI, they will also train them on other, diverse tasks. So, you could say that they also only improve them on the tasks they trained on, but they train on everything. It's like a joke when people complain that large language models (LLMs) only learn what's in the training data. And the other replies: " Then we'll stuff everything in there." So we had these operations, Python operations. For example, increment, doubling, parity. These are just a few Python tasks , and the model improved the results in them. Now, it's interesting that she didn't improve on this task. I don't know why. I might have to check that out. So, maybe it's a bug in my code, or an AI-generated bug, or an error in my experiment. Or if it's not a mistake, which is even more interesting, but it's most likely a mistake. And on the tasks she didn't practice on, she didn't improve. Perhaps it even degraded. Now, depending on how we calculated statistical significance, one should always be cautious. Has it really degraded? Do you have enough samples? Did you calculate everything correctly? It may be worth running the sample several times to see the average at different percentages. Consequently, two previously strong groups degraded because only a narrow parameter plane was optimized for reward distribution. But she improved in what we trained her on in our environment. This means you can teach her anything if you know how to build environments. So during experiments, it's a good idea to change only one variable at a time. Let's say you have these different variables: reward, reward method, group size, number of responses generated at once (e.g., eight per query), temperature as a randomness or penalty, negative rewards. Therefore, if you are conducting research, you need to do one experiment, one variable at a time. Here I conducted a number of different experiments. You see, if I set, for example, a group size of four, it works very well, better than with a group size of 16. Also, the temperature of 0.2 is more deterministic , less random, and works better than the higher random temperature. And then , uh, here in these experiments on binary rewards, like 0-1 or other fractional rewards, or no penalty for invalidity, a hard penalty for invalidity— they all work the same on my dataset. Now, you don't want to have just five, you want to have about 500 or 5,000 of those to get better insights. But look at this horror . You see, a group of four beat a group of eight on this seat, but lost on this one and lost on that one. Okay, so that means, first of all, it doesn't have much of an impact based on my settings and my training data. Second, yes, we need to train more to see true statistical significance. Um, and it's also possible that it's not that important. But if you are training huge models, even small differences can mean significant cost savings. Your model will train a little faster, which will save you a lot of money. So when people review your work, they will tell you to proofread in three sittings. And all three sessions must give a clear conclusion. For example, all three seats should indicate that group four is better than group eight. So, none of my experiments produced statistically significant improvements. Maybe my experiments are too small. Perhaps as you train larger models, you will have a clearer and clearer understanding of exactly which parameters will be best. And here we can consider each part. So, the previous training will teach her to write in Python. The revision will teach her to add image recognition to the transformer, to the language model. And reinforcement learning in our case will improve the increment and doubling operations. So I encourage you to just go and talk to Codex or cloud code. You can ask him, "What experiments can I do here?" And don't worry so much about scaling. I just want you to learn how to conduct experiments, learn scientific thinking, how to measure correctly and draw conclusions. So, you can publish your experiments on social media as well. If you want to become an AI researcher, I recommend that you continue to post your experiments daily or weekly on social media. Consistency is more important than intensity. So if you post for five days in a row and then stop—that's not good. If you post once a week consistently, that's much better. And it 's better to just conduct experiments and train the model for a few seconds. You don't want to wait a few minutes or hours because that's a waste of time. You need to think about how to learn to design experiments, environments, questions. You don't need to wait for a model because you won't be able to train a model that will outperform these companies that people will use.
Generated algorithmically for Search Engine
Indexing.