Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

How to build function calling and JSON mode for open-source and fine-tuned LLMs

Blog post from Baseten

Post Details
Company
Date Published
Author
Bryce Dubayah, Philip Kiely
Word Count
1,339
Company Posts That Month
4
Language
English
Hacker News Points
1
Post removed?
No
Summary

NVIDIA has announced support for function calling and structured output for LLMs deployed with its TensorRT-LLM Engine Builder, adding model server level support for two key features. Function calling allows users to pass a set of defined tools to an LLM as part of the request body, while structured output enforces an output schema defined as part of the LLM input. These features are built into NVIDIA's customized version of Triton inference server and use logit biasing to ensure valid tokens are generated during LLM inference. The implementation has minimal latency impact after the first call with a given schema is completed, allowing for efficient use of these new features.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 25 3,889 441 129 +7%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.