Repository navigation
Make it easier to traverse the frame stack for third party tools. #100987
Description
Activity
Initially, I propose to refactor the PyInterpreterFrame struct such that it starts:
typedef struct _PyInterpreterFrame {
PyCodeObject *f_code;
struct _PyInterpreterFrame *previous;
...
Currently f_code must be a code object, but we could generalize it to allow other objects.
For example, the shim frame inserted on entry to _PyEval_EvalFrameDefault could have that field set to None indicating it should be skipped in tracebacks, etc.
The order of f_code and previous doesn't really matter, but have f_code first makes #100719 a bit simpler
Let me collect some feedback from maintainers of debuggers and profilers and will comment here the requirements so we can think of solutions.
@pablogsal Any feedback?
We can further improve traversal of the _PyInterpreterFrame for debugging and introspection by allowing C extensions to create frames without the rigmarole of creating a code object.
We should rename the f_code field to f_executable, and allow any object.
typedef struct _PyVMFrame {
PyObject *f_executable;
struct _PyVMFrame *previous;
} PyVMFrame;Although tools and the VM should tolerate any object, we should in practice only allow a few classes:
- CodeObject: Implies that the
PyVMFrameis a full_PyInterpreterFrame. Only the VM should make this kind of frame - Builtin function, method descriptor, slot wrapper, etc. The frame represents a call to the given object.
None: Internal shim. Tools should skip this frame.- Tuple: First three items should be
name,filename,flagswhereflagsdetermine the meaning of additional entries.
The tuple form is for tools like Cython, Nanobind, etc. Creating a tuple of strs and ints is much simpler and faster than creating a fake code object.
C extension can link themselves into the frame stack at the cost of about 4 memory writes, and 3 reads:
PyVMFrame frame;
frame.previous = tstate->current_frame.frame;
frame.f_executable = &EXECUTABLE_OBJECT;
tstate->current_frame.frame = &frame;
/* body of function goes here */
tstate->current_frame.frame = frame.previous;
We can do this for builtins functions by modifying the vectorcall function assigned to the builtin function/method descriptor.
We would need to benchmark this to see the performance impact, but it will be much cheaper than sys.activate_stack_trampoline()
@pablogsal Any feedback?
I have reached out again to tool authors, give me a couple of days to gather comments. Apologies for the delay
No problem.
@benfred -^
I've made a branch that adds "lightweight" frames (just a pointer to a "code" object and a link pointer), and inserts one for each call to a builtin function in the interpreter. The performance impact is negligible and all builtin function and class calls are present in the frame stack.
I've made a branch that adds "lightweight" frames (just a pointer to a "code" object and a link pointer), and inserts one for each call to a builtin function in the interpreter. The performance impact is negligible and all builtin function and class calls are present in the frame stack.
We still need the concept of entry frames for tools that merge native and python stacks. Why do you removed _PyFrame_IsEntryFrame in your branch?
30 remaining items
How do you get the frame without any symbols?
Find the interpreter state and having the headers vendor so we know the offsets to the pointers in every struct and we know what we are going to find because at the moment is fully determined. The interpreter state can be found because we (cpython) place the runtime structure in its own section so it can be found without symbols:
Line 100 in 7703def
| __attribute__ ((section (".PyRuntime"))) |
Although this is technically not needed because it can be found by finding the cycle interpreter state <-> thread state by scanning the bss which is what py-spy does.
I don't see the value in tagging bits. The tag you propose holds no additional information, as the same value can be got with the simple comparison
f_executable == &PyCodeObject
Sure, if you have &PyCodeObject available.
Tagging bits could "make life easier for a few tool authors" in the scenarios @pablogsal is mentioning without "making things slower for very many Python users."
EDIT: also, it's not f_executable == &PyCodeObject, it's f_executable->ob_type == &PyCodeObject, so it's adding an extra pointer chase for every frame also.
Tagging might solve the performance issue. But we need to support 32bit machines, so we only have 2 bits to play with, which is not enough.
If tools can find the runtime, then we can put an array of pointers there. No runtime overhead at the cost of ~40 bytes.
PyObject *callable_types[] = {
&PyCode_Type,
...
};If tools can find the runtime, then we can put an array of pointers there.
That would be an acceptable compromise I think.
OK, let's go with that then.
FTR, one other reason not to use an enum is this: what happens when the enumeration and the executable don't match?
We can be fairly sure it won't happen in our code, but it would be an easy mistake to make in third-party code.
By allowing objects of any class, but designating a small set of "approved" classes, the system is much more robust.
@pablogsal
Where should the array go, exactly?
I hope we can make life easier for existing inspection tools by making it really easy to detect the common cases they want to care about, but I also hope (from the Cinder JIT perspective) that at least one of the valid options for f_executable is "extensible". E.g. if(name, filename, lineo) tuple is allowed, that it's also valid to have a longer tuple carrying additional payload, with the first three elements interpreted as name, filename, lineo.
I don't see a problem with that.
Tools should check the length of the tuple before extracting the contents, for safety.
We could allow any length array, specifying only that the first three elements, if they exist, should be the name, filename and line number.
("foo",) and ("foo", "foo.py", 121, "special-data-34.8") should both be acceptable.
@pablogsal Would this be OK, or is this too complex for your tastes?
Some comments from authors:
I feel it won't be too easy to decipher the type of the object remotely. This would likely increase the number of private structures that we need to copy over from Python headers to parse this information (e.g. tuples), making things more complex. Of course one could just try treating the object as a PyCodeObject and check for failures, but this would now imply a potential loss of captured information, unless all the other object types that can appear here are also handled. Perhaps an extra int field that specifies the type of the object being passed with f_executable might help in this direction, to some extent. But perhaps one simplification that depends on a positive answer to the following question could be adopted: is the value f_executable crucial for the actual execution, or is it just added to carry the frame's metadata (e.g. filename, function name, line number, ...)? If that is added just for the metadata, perhaps that could be added directly to the _PyVMFrame structure in the form of extra fields? There could be a core set of fields that are common to all object types (filename, function qualname, location data), plus a generic PyObject reference that can be consumed easily by in-process tools. However, I can see the downside being that the cost would probably end up being slightly more than just 4 W and 3 R operations in general.
This would be me, maintainer of Austin. For context, Austin uses system calls like process_vm_readv to read memory out of process.
How do you get the frame without any symbols?
Austin uses symbols to locate _PyRuntime, but if those are not available, there is a fallback on BSS scan to locate something that looks like _PyRuntime or an interpreter state. So symbols are not strictly required (but good to have of course!).
Apologies if I slightly derail the conversation, but I wanted to express the following thought. Based on my experience with Austin, I would regard frame stack unwinding as just one aspect of the more general topic of observability into the Python VM. For example, one other thing that Austin tries to do is to sample the GC state to give an idea of how much CPU time is being spent on GC. Or detect who is holding the GIL to give a better estimate of RSS allocations. Therefore, I would tend to view frame stacks as just a part of what can be observed out of process. So the way I see a tool like Austin extracting this information in the future is by looking into an "observability entry point", much like _PyRuntime, but specifically engineered for out-of-process tools. From there one can rely on an ever growing (in an ideally backwards-compatible fashion) list of things one can observe, e.g.
.section _PyRuntimeStateABI
runtime_state {
interpreter_state {
thread_count: int,
threads: [{
top_frame: { ... },
....,
]
},
gc_state: ...,
gil_state: ...,
}
Profilers and debuggers need to traverse the frame stack, but the layout of the stack is an internal implementation detail.
However can make some limited promises to make porting tools between Python versions a bit easier.
In order to traverse the stack, the offset of the
previouspointer needs to be known. To understand the frame, more information is needed.@pablogsal
@Yhg1s
expressed interest in this.
Linked PRs
_PyInterpreterFramea bit, to assist generator improvement. #100988_PyEval_EvalFrameDefault. #102640