· Originally published on Medium

LLMs, the good and bad, and the trend

This is an article I wrote last night. With the roll out of ChatGPT Plugins, the game has changing yet again. But many points in the article still make its point. But everything…

  • ai
  • openai
  • chatgpt

Photo by ilgmyzin on Unsplash

Photo by ilgmyzin on Unsplash

This is an article I wrote last night. With the roll out of ChatGPT Plugins, the game has changing yet again. But many points in the article still make its point. But everything is going to happen sooner than we can imagine.

What is…

GPT stands for Generative Pre-trained Transformer. It is an autoregressive language model that uses deep learning to produce human-like text. Despite the technical jargon, you don’t need to understand it to continue reading.

In my understanding, GPT is a powerful method of lossy compression for pure-text information. Despite the extensive training required for these models, their pre-trained size is surprisingly small. This represents a significant advancement for information science, allowing us to permanently increase the information density possible in human history.

While the human-like interaction may be deceiving, there is no denying that conversation is the most effective method for humans to gain access to information. And that is how the model is designed to work most efficiently: token in, token out.

From a neural network perspective, GPTs are not like humans. Humans are sophisticated parallel computing machines based on a tree-stack of neural networks, triggered by time and inputs from many sources. GPTs, on the other hand, are much simpler. Therefore, the “intelligence” they exhibit is merely the behavior of the deeply compressed, lossy information being handled delicately by the network itself. With the rollout of GPT-4, however, the model becomes increasingly “clever,” which may be another example of how quantitative change produces qualitative change. With continued advancements in deep learning and neural networks, we may gain a better understanding of how consciousness is formed.

How is…

From a practical standpoint, GPTs are useful. They are essentially know-it-all text-processing machines that outperform the search engines we know. However, more importantly, they are better at text processing than information search. As we discussed earlier, GPTs are simply algorithms that compress knowledge with loss. This means they cannot identify the truthfulness of information. They can only access the stored information from the model and try to process or restore the compressed data in the form of human language. The ability to process human language is the real crown jewel of this type of model.

This is where hallucination comes in. I won’t delve into the technical details of hallucination, but you can think of it like attending an exam without proper preparation. Some of the information is not stored in your brain, but you can force yourself to fill in related memories with a bunch of nonsense. Even if it’s a math exam, you can try to relate something on the paper. Hallucination is similar to this.

AI doesn’t have the ability to discern what is right and what is wrong, so the results could be “confidently wrong.” Therefore, for now, I suggest only using the information query feature with DYOR (Do Your Own Research) in mind. It cannot be trusted under any circumstances.

How to…

There’s not a lot of ways for us to integrate GPTs into our daily work. Most of us use different tools everyday to improve our efficiency. Naturally, the most efficient way to integrate AI into our daily work is to wait the tools we use integrate AI into them.

But there’s two worse case. One is the tools won’t integrate AI no matter what reasons, not likely, but still possible. Another is, even integrated, it is not something we want.

But once they are integrated for particular products, they will have a huge leap on co-working experience, because they have already been fine-tuned

Now, the most efficient way to interact with GPTs are using Chat-like UIs. Like ChatGPT or other alternatives.

Prompt Engineering

Prompt engineering isn’t something that literally involves engineering. When we are giving input to the ChatGPT, we are adding more constrains to the answer.

ChatGPT have a limited amount of “memory”, that’s because it is not actually “memory” like we have. The reason ChatGPT and relate with context is because the UI will send the chat history along with the new prompt. So, the more we talk, the more context will send to the ChatGPT at one prompt.

The easiest example of prompt engineering is to tell the model the thing you want it to do, then inside the thread(session) you created, it will do the thing you ask it to. If the context is within the input limit.

For example. I built a thread with OpenCat, which is a GUI for OpenAI API. I create a new thread with a pre-embedded prompt like: “Translate each sentence I send to you to French. The text will be use on UI, so be lean and short.”

Than, every sentence I input into the thread will get a response from ChatGPT with short and lean French translation I want.

Prompt Engineering is good at things with simple input. Like, role play or discussion within topic, or translation with preference. Because GPT doesn’t have logic ability, so it is very poor on math. And it is better to keep the thread as short as possible, this can prevent hallucination.

Fine-tuning and Embedding

These are the true production-ready solution for future GPT models. Let me explain them one-by-one.

Fine-tuning and embedding are two important concepts for future GPT models. Fine-tuning involves customizing a pre-trained model for a specific task or domain. This allows for better performance on specific tasks and can reduce the amount of training data needed. Embedding, on the other hand, involves mapping words or phrases to numerical vectors in a high-dimensional space. This allows for easier comparison and manipulation of text data. As GPT models continue to advance, fine-tuning and embedding will play increasingly important roles in their development and implementation.

Hard to understand? It’s OK. Let’s make some examples.

Suppose we aim to train a GPT model for generating recipe-related text. First, we must fine-tune the model using a dataset comprising recipe texts. Fine-tuning entails customizing the pre-trained GPT model for the specialized task of recipe generation by extending its training with new data, in this case, recipes.

Upon fine-tuning, we can utilize the model to create novel recipes by giving it a prompt or context (e.g., “Ingredients: flour, sugar, eggs”) and requesting the completion of the recipe.

Nonetheless, for the GPT model to produce precise and coherent recipes, it must comprehend the significance of individual words and their interrelationships within a sentence. This is where embedding plays a role.

Embedding refers to the representation of words in a manner that encapsulates their associations with other words in a sentence or dataset. In our recipe scenario, embedding would involve representing each ingredient and cooking instruction such that the GPT model can learn to generate new recipes based on this information.

Thus, fine-tuning and embedding are two synergistic techniques employed in training GPT models (and other deep learning models) for specific tasks like text generation. Fine-tuning allows the model to adjust to a particular dataset, while embedding enables the model to grasp the meanings and relationships of individual words within that dataset.

Fine-tuning for learning, embedding for classification.

If you want to learn more about how, here is a more detailed technical document.

Ok, But…

The limitation

Although GPT technology has shown promise, it is still an immature field with numerous limitations. As a result, the future lies in multi-modal GPT models that incorporate various data sources, such as images, videos, and other constructs. Humans perceive the world through real-time, high-quality video and audio feeds, with language serving as an abstraction of spoken communication. Unfortunately, some information is inevitably lost in this process.

To address this issue, we can enhance the emotional impact and emphasize certain aspects of text by incorporating “metadata” like emojis 😄, **markdowns**, or comments. However, these are currently the only forms of metadata that GPT models can accept. There is still much progress to be made in terms of accuracy, safety, and overall capability of these models.

The price

While ChatGPT offers numerous benefits, it comes with a significant financial cost. To use GPT daily, users must spend $20 per month, invest time in learning the system, and dedicate substantial effort to crafting prompts and refining the generated results. All of these factors contribute to the time-consuming nature of the process. Ideally, GPT should function as a supportive tool for individuals, offering potential solutions rather than being a direct replacement for human input.

The effort

The migration to GPT technology will require significant effort and investment on the part of organizations looking to integrate it into their workflow. This includes not only the financial cost of using GPT, but also the time and resources required to train and fine-tune the model for specific tasks. Additionally, there is a risk of overreliance on GPT and a loss of human input and creativity in the decision-making process. Therefore, organizations must approach the adoption of GPT technology with caution and a clear understanding of its limitations and potential pitfalls.

Natural Language Programming

GPT technology, despite its advantages and drawbacks, heralds a new era in Natural Language Processing, offering immense potential for the field. With GPT, particularly in conversational applications, users can engage with a vast repository of knowledge akin to the entire internet, using everyday human language as their programming tool.

Although current limitations will gradually be overcome, the emergence of GPT signifies the dawn of a “META” age, where more powerful capabilities and fewer restrictions will pave the way for even more advanced natural language programming.

All articles