2.7 Creating token embeddings

This content explores the fundamental preprocessing steps required for preparing text for large language models, focusing on how to tokenize input into individual words and special characters. It demonstrates practical methods for text splitting using Python's regular expressions and explains the subsequent conversion of these tokens into unique numerical IDs by constructing a vocabulary, which is an essential precursor to creating embedding vectors for machine learning.

This page is free — overview and chapter list only.
The complete book body is sold separately.

Full book: $2.99 USDC via x402 ·
Per chapter: $0.25 USDC

Buy / open complete book (HTML)
· Complete book Markdown (.md)

Agents: start with free /library/discovery.json, sample free teaser chapters,
then pay for individual chapters or the complete book URL above.
Append .md to any content URL for Markdown with YAML front matter.

Chapters

  1. 22 Tokenizing Text (Free teaser)
  2. The Output Is ($0.25)
  3. 24 Adding Special Context Tokens ($0.25)
  4. The Output Is 2 ($0.25)
  5. 25 Byte Pair Encoding ($0.25)
  6. 26 Data Sampling With A Sliding Window ($0.25)
  7. The Code Prints ($0.25)
  8. The Print Function Call Returns ($0.25)
  9. This Chapter Covers ($0.25)
  10. 31 The Problem With Modeling Long Sequences ($0.25)
  11. The Self In Self Attention ($0.25)
  12. 331 A Simple Self Attention Mechanism Without Trainable Weig ($0.25)
  13. Exercise 33 Initializing Gpt 2 Size Attention Modules ($0.25)
  14. The Output Is 3 ($0.25)
  15. The Encoded Ids Are ($0.25)
  16. 513 Calculating The Training And Validation Set Losses ($0.25)
  17. The Cost Of Pretraining Llms ($0.25)
  18. Adamw ($0.25)
  19. 531 Temperature Scaling ($0.25)
  20. 532 Top K Sampling ($0.25)
  21. Listing 54 A Modified Text Generation Function With More Div ($0.25)
  22. Exercise 53 ($0.25)
  23. Exercise 54 ($0.25)
  24. The Contents Are ($0.25)
  25. Exercise 56 ($0.25)
  26. This Chapter Covers 2 ($0.25)
  27. Choosing The Right Approach ($0.25)
  28. Exercise 61 Increasing The Context Length ($0.25)
  29. Listing 66 Loading A Pretrained Gpt Model ($0.25)
  30. The Model Output Is ($0.25)
  31. Output Layer Nodes ($0.25)
  32. Fine Tuning Selected Layers Vs All Layers ($0.25)
  33. Exercise 62 Fine Tuning The Whole Model ($0.25)
  34. The Initial Loss Values Are ($0.25)
  35. Choosing The Number Of Epochs ($0.25)
  36. The Resulting Accuracy Values Are ($0.25)
  37. Exercise 71 Changing Prompt Styles ($0.25)
  38. Listing 74 Implementing An Instruction Dataset Class ($0.25)
  39. Exercise 72 Instruction And Input Masking ($0.25)
  40. A13 Installing Pytorch ($0.25)
  41. Pytorch On Apple Silicon ($0.25)
  42. This Prints ($0.25)
  43. A23 Common Pytorch Tensor Operations ($0.25)
  44. The Output Is 4 ($0.25)
  45. The Output Is 5 ($0.25)
  46. Listing A2 A Logistic Regression Forward Pass ($0.25)
  47. Listing A3 Computing Gradients Via Autograd ($0.25)
  48. A5 Implementing Multilayer Neural Networks ($0.25)
  49. This Prints 2 ($0.25)
  50. This Prints 3 ($0.25)
  51. The Result Is ($0.25)
  52. The Result Is 2 ($0.25)
  53. The Result Is 3 ($0.25)
  54. The Result Is 4 ($0.25)
  55. Running This Code Yields The Following Outputs ($0.25)
  56. Exercise A3 ($0.25)
  57. This Outputs ($0.25)
  58. The Output Is 6 ($0.25)
  59. A91 Pytorch Computations On Gpu Devices ($0.25)
  60. Listing A11 A Training Loop On A Gpu ($0.25)
  61. A93 Training With Multiple Gpus ($0.25)
  62. Selecting Available Gpus On A Multi Gpu Machine ($0.25)
  63. Summary ($0.25)
  64. Chapter 1 ($0.25)
  65. Chapter 4 ($0.25)
  66. Appendix A ($0.25)
  67. Exercise 22 ($0.25)
  68. Exercise 41 ($0.25)
  69. Exercise 43 ($0.25)
  70. Exercise 53 2 ($0.25)
  71. Exercise 63 ($0.25)
  72. Exercise 74 ($0.25)
  73. Exercise A2 ($0.25)
  74. When Executed On A Gpu ($0.25)
  75. E4 Parameter Efficient Fine Tuning With Lora ($0.25)
  76. Symbols ($0.25)
  77. Related Manning Titles ($0.25)
  78. Hands On Projects For Learning Your Way ($0.25)
  79. Explore Dozens Of Data Development And Cloud Engineering Liv ($0.25)
  80. Whats Inside ($0.25)