Monday, August 29, 2022

JIT Compiler Dongle - The Connection HPC 2022 RS

JIT Compiler Dongle - The Connection HPC 2022 RS (c)Rupert S


JIT Compiler Dongle makes 100% Sense & since it has no problem acting like a printer! It can in fact interface with all printers & offload Tasks,

However in High Performance Computing mode of operation the USB Dongle acts as the central processor from the device side; That is to say the device such as the printer or the Display...

You can supply a full workload to the dongle & of course it will complete the task with no necessity of assistance from the computer or the device.

The JIT Compiler comes into its own one two fronts:

Compatibility between processor types.

Aiding a device in processing &or passing work to that device to run; Work that is shared & if required workloads are passed back & forth & shared,

Shared & optimised...

The final results for example are post-scripts? no problem!
The final results for example are Directly Compute Optimised Printer Jet algorithms? no problem!
The task needs to compute specifics for a DisplayPort LED Layout ? no problem!

The device is powerful so share, JIT Compiler for real offloading & task management & runtime.

Functional Processing Dongle Classification USB3.1+ & HDMI & DisplayPort (c)RS

Theory 1 Printer

Itinerary:

Printers of a good design but low manufacturing cost of ICB printed circuits have a printhead controller,

But no Postscript Processor; But they do have a print dither controller & programmable version need to interface with the CPU on the printing device,

Print controlling is a viable Dongle & also Cache but workload cache has to have a reason!

That reason here given is the JIT Dongle that is able to interface with both Web print protocol & IDF Printing firmware.

But here we have postscript input into the JIT Compiles Kernel & output in terms of Jet Vectors & line by line Bitmap HDR & head motion calculations,

We can also tick the box on Postscript offloading on functioning PostScript printers; But we prefer to offload JIT for speed & size..

Vectors & curves & lines & Cache.

Theory 2 Screen

Itinerary as of printers but also VESA & line by line screen print & VESA Vectors & DisplayPort Active displays,

Cable Active displays require the GPU to draw the screen & calculate the Line Draw!

The Dongle activates like a screen with processor & carries the screen processing out; Instead of a smartwatch or small phone that does not have a good capacity for computer lead active display enhancements.

Theory 3 Hard Drives & controller such as network cards & plugs for PCI

Adapting to Caching & processing Storage or network data throughput commands, While at the same time being functionally responsive to system command & update makes JIT Dongle stand out at the head of both speed & function...

Network cards can send offloading tasks to the PCI socket & the plug will process them.

Hard-drives can request processing & it shall be done.

Motherboard ROMs & hardware can request IO & DMA Translation & all code install is done by the OS & Bios/Firmware.

Offloading can happen from socket to Motherboard & USB Socket & URT..

All is done & adapts to Job & function in host.

The 8M Motherboard & OS verifies the dongle, licences the dongle from the user..
& runs commands! Any Chipset, Any maker & every dongle by Firmware/Bios
What the unit constitutes is a functional Task offloader for OS & Bios/Firmware.

The utility is eternal & the functions creative & secure & licensed/Certificate verified.

Any Motherboard can be improved with the right Firmware & Plugin /+ device.

(c)RS

*****

DDM Super Immediate Display Modes with 0ms GTG : Operation Latency Zero


By initiating DDM & using the display processor aswell with DPIC JIT,
With DDM Frame Buffer Emulation & Control.

Games & Aiming for Business,
DDC & FreeSync Update today!

In order to set DDM Super Immediate Display Modes you have to set the
display as being DDM with an input frame buffer..

That way both the GPU & the Display can work on the frame in ALLM
Mode; Enhancing processing while reducing latency.

*
FreeSync - DDM - Low Latency Screen Modes

HDMI & DisplayPort : Screen Framebuffer {DDM, FreeSync, ALLM} : Minimal latency post processing : RS

Direct Drive Monitor (DDM) is a mode where the Frame is directly created by the CPU/GPU facing the screen,
The frame buffer facing the Processor must present all capacities & properties of the Screen directly..

List of common properties:

Frame Buffer & Frame buffer write control
Bit depth & FRC
DSC mode
ICC Colour Profile
Write Cache buss width
Timings
Latency
LED Colour range & profile

The GPU/CPU must have the capacity to order write cycles & DSC Decompression layer,
The GPU/CPU must not have a discussion writing to the screen; Direct Write shall be immediate!

So we need to have the frame buffer process as fast as possible & report back,
But we plan to initiate a frame buffer & process it!; Process the frame fast,
To do that we provide all the information from our frame buffer that the CPU/GPU needs to calculate..

Rupert S
*

A DDM Monitor is directly controlled by a GPU/CPU


Initiating a Direct Drive Monitor (DDM) capability enables ultra-thin monitors (and Mobile Phone Screens) With a Short Plug DSC Compression Array...

Could be simple!

Initiate a DDM Mode with DPIC: JIT Kernel to a Frame Processing Unit that directly presents as a surface; All tasks from there in will not be allowed to add latency.

(DDM) DPIC JIT Compiler Mode handles the situation of under performing hardware quite well,

The aim is to solve one of the largest issues with DDM & that is latency! & Frame Distortion such as Frame Blur,

Long cable access to a device encounters the same latency issues as RAM & Storage,
Distance means time!

By Directly compiling commands into an (ESK) Efficient Static Kernel; Stack space (Cache & RAM)...

Processing load is light & may be performed On The Edge; Close to the hardware; in our case a screen with a Single Core ARM Nano millimeters close to the screen.

No we do not need a large CPU that close; But a SiMD array & Texture decompressor & Direct screen print...

We do all our Large Problem solving previously in JIT Kernels; While doing what we can closer to the screen at our Frame Buffer,

We can also directly process commands directed from a larger processor; a CPU, GPU, HUB,

All we need to do is Initiate a DDM Mode with DPIC: JIT Kernel to a Frame Processing Unit that directly presents as a surface; All tasks from there in will not be allowed to add latency.

Rupert S

*

Direct Drive Class : Displays, Printers & Devices such as Joysticks, Mice & keyboards


You know Active Display,
The DisplayPort & HDMI Configuration,
JIT Compiler is a way of getting these to work internally inside the GPU &
In Port class units & USB Dongles that process Computation tasks,

The JIT Compiler DPIC System processes for the Display,

Therefore Able to Activate the display to the highest level of
processing with minimal requirements of necessity!

For example Active Displays with basically a Micro NUC that has an arm
processor & is 4 CM² with USB Connection,
Therefore can power an active display (the type with smaller processors)

Additionally can carry out more work & share a single NUC with
multiple Active Displays..

Bearing in mind that such a OpenCL/JIT Driver is universal to all Systems & Simply classifies by processor
class.

*

The primary motivation for Direct Drive Class displays & Equipment is to offload Processing tasks to the GPU/CPU...


However by example we can Flow Control frames on the HDMI & DisplayPort cables,
We do this by Writing a Kernel/OpenCL Code (Around 60KB) that queries the Frame Ready Flag/Property in the GPU...

Example of Coding Model {Display CPU <> GPU} : Audio : Video : Texture Set

OpenCL Kernel Runtime 512Kb (aim)

Set Properties of display screen (Size & compression & Unique properties such as Texture Types)
Request Frame memory Allocation
Frame Pull (Demand a frame)
Query Frame & Send Ready flag

When Frame Ready***

Send workloads to GPU on frame: Example
Decompression Stack
Frame Mask
Memory Load (Direct DMA access to RAM from Cable)

Sort functions,
Optimisation tasks such as Colour range optimisation & WCG, HDR Tone Mapping.

We keep these operations from sending frame & texture back; by operating on the frame analysis before sending frame...

Reception process involves sending:

Data From Tasks first
RAM Page Map (if we did this process on GPU)
Frame
Process (We send additional tasks if required from a worker thread)

Send out Query : Repeat!

#GoodFramingDirectDrive

RS

***

Plan 2023-03-07 Direct Map DDM : Efficient Monitor Direct Frame Forwarding Render : DDM ALLM : Rupert S


DDM Combined with Combining Texture converters, FSR & OpenCL (Compilable for processor types & Firmware),

Allows the HDMI & DisplayPort FrameBuffer Abstraction layer to pass fully optimized texture layers directly into Frame Rendering & therefor to be directly DMA Copied to the screen along with the Colour conversion table mapping (can be done by the GPU, The Monitor or be Hardware intrinsic to DSC.

DPIC JIT Compiler (Kernels : Small & into Precompiled Code Array Buffer)

https://drive.google.com/file/d/1D27MOBYKVkKib1JzP_eFucp8RRrzAhd6/view?usp=sharing, https://drive.google.com/file/d/1DbcifAxrG4XKfJ9Mrpsfq7kq1I4aV5ES/view?usp=sharing, https://drive.google.com/file/d/1d_bWbZl9fAZXsLbN_jZdqSxdWzraLSIz/view?usp=sharing
***

GTG : GoodToGame : Consoles, Gaming & Movies : RS

DDM ALLM Dongle use case - 10K Presentation of abstract data polygons

Dear VESA & HDMI; This gaming feature does rely on the GPU (in the main) but does have CPU Capacity; Particularly in consoles of the new generation!

Luckily in my experience the CPU is often under utilized by Vulkan API & DirectX 12.1 & therefore the use of DDM mode combined with the JIT Compiler OpenCL compile is a very logical choice! due to DDM being VESA we arrange it,

Clearly the RAMDAC does pre compute a frame & clearly a frame is no more than 15% work for a FreeSync monitor,

Combining the strengths of 2016+ TV & monitor ARM processors (600Mhz+; obviously less powerful than 550Mhz & DDM + ALLM + FreeSync + JIT Compiler is a clean logical choice)

RAMDACS are 600Mhz but we can multitask; So we shall.

Rupert S on behalf of VESA & HDMI & the gaming & Film community.

*****

Example Display Chain (Can be USB/Device Also For the OpenCL Runtime; To Run or be RUN) (c)RS


How a monitor ends up with an OpenCL : CPU/GPU Run Time Process: Interpolation & Screen enhancement: The process path

Firstly we need to access the GPU & CPU OpenCL Runtime such as:

Components that we need:

https://science.n-helix.com/2022/08/jit-dongle.html

https://science.n-helix.com/2022/06/jit-compiler.html

https://science.n-helix.com/2022/10/ml.html

FPGA 'Xilinx Virtex-II' HPC application Multiple-Applications & Image-Net & Matrix-Multiplication - H-SIMD machine _ configurable parallel computing for data-intensive HPC
https://digitalcommons.njit.edu/cgi/viewcontent.cgi?article=1836&context=dissertations

A SIMD architecture for hard real-time systems
https://www.repository.cam.ac.uk/bitstream/handle/1810/315712/dissertation.pdf?sequence=2

Ideal for 4Bit Int4 XBox & Int8 GPU
PULP-NN: accelerating quantized neural networks on parallel ultra-low-power RISC-V processors - Bus-width 8-bit, 4-bit, 2-bit and 1-bit
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6939244/

https://science.n-helix.com/2023/02/smart-compression.html

Firstly, we need an OpenCL Kernel : PocCL :

PoCL Source & Code
https://is.gd/LEDSource

MS-OpenCL
https://is.gd/MS_OpenCL

HIP_CUDA on OpenCL & SPIRV
https://github.com/CHIP-SPV/chipStar

HIP_CUDA on OpenCL & SPIRV : ZIP
https://is.gd/HIP_CUDAonOpenCL

X86Features-Emu
https://drive.google.com/file/d/15vXBPLaU9W4ul7lmHZsw1dwVPe3lo-jK/view?usp=usp=sharing

Crucial components:

Microsoft OpenCL APP
Microsoft basic display driver OpenCL component (CPU)

CPU/GPU OpenCL Driver
PoCL Compiled runtime to run Kernels https://is.gd/LEDSource

We need an Ethernet connection to the GPU (Direct though the HDMI, DisplayPort),
A direct connection means no PCI Bus or OS Component needed,
(But indirect GPU Loaded OpenCL Kernel loading may be required)

Or

We need an Ethernet connection to the PC or computer or console!
Then we need a Driver (this can be integral or Drive) to load the OpenCL Kernel; This can have 3 parts in the main to run it!

Microsoft OpenCL APP
Microsoft basic display driver OpenCL component (CPU)

CPU/GPU OpenCL Driver
PoCL Compiled runtime to run Kernels https://is.gd/LEDSource

The compiled Kernel itself & this can be JIT : Just In Time Compile Runtime

Rupert S

*****

Example of low latency Kernel Runtime to run process sharing between the processors of a system.

VESA DDM is much better with a resident kernel in the GPU & CPU that can run the tasks of display refreshing & forward frame rendering..

One simply needs to know that the display cable has ethernet; Through Ethernet local networking the OpenCL Kernels can be compiled & run..

Enabling any tasks to be shared & averaging processor load & usage of an entire system.

RS

*****

Offloading JITcompiler : Device Driver Offloading Acceleration Pool


Microsoft makes a note of Audio Offloading in the latest chrome update,

It makes sense to note that JITcompiler is a highly efficient transparent offloader,

With CPU & GPU & Processor OpenCL compilers; There is a good chance of your function being supported!

Simple or complex maths & logic is ideal for OpenCL & the creation of a general pool for offloading functions..

OpenCL is able to cater to many maths functions & with calls to security & networking functions,

OpenCL will be very capable in that capacity.

Rupert S

References : https://is.gd/DictionarySortJS

https://science.n-helix.com/2023/06/tops.html
https://science.n-helix.com/2022/08/jit-dongle.html
https://science.n-helix.com/2022/06/jit-compiler.html

https://science.n-helix.com/2021/03/brain-bit-precision-int32-fp32-int16.html
https://science.n-helix.com/2022/10/ml.html

Firstly, we need an OpenCL Kernel : PocCL :

PoCL Source & Code
https://is.gd/LEDSource

MS-OpenCL
https://is.gd/MS_OpenCL
https://is.gd/OpenCL4X64
https://is.gd/OpenCL4ARM

*****

The DPIC Protocol in use for display, robotic hardware (arms for example) & Doctor Equipment arms & surgeries, Website loading or games.


In context of load for DPIC, We simply need a page (non-displaying Or Displaying (for example Monitor Preferences)) Inside the GPU..

Can use WebJS, WebASM : WASM, OpenCL : WebGPU : WebCL : WebGPU-ComputeShaders...

RAM Ecology wise between 1MB to 128MB RAM (But should inform client in print of options); I cannot really imagine you would need more apart from complex commands (cleaning for example & robots)

Direct Displayport & HDMI Interface; With or without use of USB Protocol HUB..

Touch screen operation examples:

Can additionally Smart pick diagnostic process of operations or equipment placement & screw & nut & bolting operations & welding or cutting!

For example, the DPIC Protocol can interface & runtime check Operations, Rotations, Motions & activations in well managed automatons; While directly interfacing the ARM/X64/RISC Processor tools & where necessary optimise memory & instruction ASM Runtime Kernel.

*

How does PTP Donation Compute work in business then:

Main JS Worker cache (couple of MB)

{ main . js }

{

{ Priority Static JS Files }

{ Priority Static Emotes & smilies (tiny) }

{ Priority Application JS & Static tiny lushi images (tiny) }

}
{

{ Work order sort task }

{ Sub tasks group }

{Compute Worker Thread }

}

*

(c)Rupert S

*****
Technology Demonstration https://is.gd/DongleTecDemo

Combining JIT PoCL with SiMD & Vector instruction optimisation we create a standard model of literally frame printed vectors :

VecSR that directly draws a frame to our display's highest floating point math & vector processor instructions; lowering data costs in visual presentation & printing.

(documents) JIT & OpenCL & Codec : https://is.gd/DisplaySourceCode

Include vector today *important* RS https://vesa.org/vesa-display-compression-codecs/

https://science.n-helix.com/2022/06/jit-compiler.html

https://science.n-helix.com/2022/08/jit-dongle.html

Bus Tec : https://drive.google.com/file/d/1M2ie8Jf_bNJaySNQZ5mqM1fD9SAUOQud/view?usp=sharing

Audio BT Codec

https://science.n-helix.com/2021/10/he-aacsbc-overlapping-wave-domains.html

DSC, ETC, ASTC & DTX Compression for display frames

https://science.n-helix.com/2022/09/ovccans.html

https://science.n-helix.com/2023/02/smart-compression.html


https://science.n-helix.com/2023/03/path-trace.html

*****

Good stuff for all networks nation wide, the software is certificate signed & verified
When it comes to pure security, We are grateful https://is.gd/SecurityHSM https://is.gd/WebPKI
TLS Optimised https://drive.google.com/file/d/10XL19eGjxdCGj0tK8MULKlgWhHa9_5v9/view?usp=share_link
Ethernet Security https://drive.google.com/file/d/18LNDcRSbqN7ubEzaO0pCsWaJHX68xCxf/view?usp=share_link

These are the addresses directly of some good ones; DNS & NTP & PTP 2600:c05:3010:50:47::1 2604:6600:2000:45::1 2604:6600:2000:2b::1 2a06:98c1:54::c12b 142.202.190.19 172.64.36.1 172.64.36.2 108.181.201.22 108.181.165.159

Sunday, August 14, 2022

SiMD Chiplet Fast compression & decompression (c)RS

SiMD Chiplet Fast compression & decompression (c)RS


*
Subject: SiMD Compression / Decompression chip of 2mm on side of die Chiplet (c)RS

Compression / Decompression chip of 2mm on side of die Chiplet (c)RS

Additional CPU & APU Compression / Decompression chip of 2mm to
feature on chiplet console APU's this is planned so that the Chiplet
does not require modification to the console APU,

Additionally to feature pin access Direct Discreet DMA for storage :

https://www.youtube.com/watch?v=1GvUdPn5QLg

*

Configuration of SiMD : Huffman & Compression : RS

To pack the majority of textures to 47 bit, one presumes a familiarity with Huffman codecs & the chaotic wavelets these present...

AVX256 Tasks x 4 = 64Bit
SiMD 16Bit x 2 = 32Bit / Alignment with AVX == x8
SiMD 32Bit x 2 = 64Bit / Alignment with AVX == x4

Closest to 47 = 40Bit Op x 2 (2.5Oe) | 80Bit/2 | 2 op x (1.5Oe)

So 40 Bit x2 parallel 6 Lanes

So on operation terms of precision :
32Bit Satisfies HDR,
40Bit Very much satisfies HDR,

16Bit satisfies JPG (basic)
64Bit satisfies LUT & Wide Gamut HDR Pro Rendering

*
Drill texture & image format (with contrast & depth enhancement)

https://drive.google.com/file/d/1G71Vd9d3wimVi8OkSk7Jkt6NtPB64PCG/view?usp=sharing
https://drive.google.com/file/d/1u2Qa7OVbSKIpwn24I7YDbwp2xdbjIOEo/view?usp=sharing

https://science.n-helix.com/2022/08/simd.html

Research topic RS : https://is.gd/Dot5CodecGPU https://is.gd/CodecDolby https://is.gd/CodecHDR_WCG https://is.gd/HPDigitalWavelet https://is.gd/DisplaySourceCode

*

GPU acceleration process : Huffman (c)RS


In the case of dictionary we create a cubic array: 16 parallel Integer cube, 32 SiMD,

FPU is used to compress the core elliptical curve with SVM Matrixing in 3D to 5D for files of 8Mb,FPU is inherently good versus Crystalline structure, We use the SiMD for comparative matrix & byte swap similarity.

It is always worth remembering that comparative operations are one of the most fundamental SiMD functions; But multiply, ADD & divide exist within SiMD,
Functional FPU code can always use arrays of SiMD to handle chaotic play in the field..

A main example is in Huffman's the variance of a wavelet from the main path,
Routes though main wavelet types are handled by table (on the amiga for example) &or FPU!
Micro changes make SiMD viable; In the same principle as a Hive & her ants.

Inherent expansion doubles the expected SiMD use; Ideally 2MB ram per cube
Taking advantage of a known quantity & precision we code-block by 16Bit to 128Bit segments.

Self correction allows us to Cube Huffman Decode into blocks, we parallelize blocks,
To (additionally) handle error we block the original compression.

"We also use fine-grained locking for the frequency dictionary, individually locking each key-value pair. Once the symbol codes have been determined, each symbol is replaced by its code, and all symbols; So are processed in parallel.

Decompression is inherently sequential, and hence much harder to parallelize. In this case, we take advantage of the self-synchronizing property of Huffman coding, which allows us to start at an arbitrary point"
Huffman source, Requires analysis https://github.com/catid/Zpng

https://vignan.ac.in/pgr20/20ES011.pdf
https://bestofgithub.com/repo/Better-lossless-compression-than-PNG-with-a-simpler-algorithm

ZPNG
faster than PNG and compresses better for photographic images. This compressor often takes less than 6% of the time of a PNG compressor
https://github.com/catid/Zpng
*

SiMD Chiplet Fast compression & decompression (c)RS


3 proposals


https://is.gd/BTSource

LZ77:
https://github.com/jearmoo/parallel-data-compression

The FastPFOR C++ library : Fast integer compression :
https://github.com/lemire/FastPFor

SIMDCompressionAndIntersection
C/C++ library for fast compression and intersection of lists of sorted integers using SIMD instructions : https://github.com/lemire/SIMDCompressionAndIntersection

Compressor Improvements and LZSSE2 vs LZSSE8
http://conorstokes.github.io/compression/2016/02/24/compressor-improvements-and-lzsse2-vs-lzsse8
http://conorstokes.github.io/compression/2016/02/15/an-LZ-codec-designed-for-SSE-decompression

Compression Science Docs


A General SIMD-based Approach to Accelerating Compression
Algorithms
https://arxiv.org/ftp/arxiv/papers/1502/1502.01916.pdf

SIMD Compression and the Intersection of Sorted Integers
http://boytsov.info/pubs/simdcompressionarxiv.pdf

Fast Integer Compression using SIMD Instructions
https://www.uni-mannheim.de/media/Einrichtungen/dws/Files_People/Profs/rgemulla/publications/schlegel10compression.pdf

Fast integer compression using SIMD instructions
https://www.researchgate.net/publication/220706907_Fast_integer_compression_using_SIMD_instructions

*****

The FastPFOR C++ library : Fast integer compression
Build Status Build Status Ubuntu-CI


https://jearmoo.github.io/parallel-data-compression/

GO

https://github.com/zentures/encoding

http://zhen.org/blog/benchmarking-integer-compression-in-go/

https://github.com/golang/snappy

The FastPFOR C++ library : Fast integer compression
Build Status Build Status Ubuntu-CI

What is this?

A research library with integer compression schemes. It is broadly applicable to the compression of arrays of 32-bit integers where most integers are small. The library seeks to exploit SIMD instructions (SSE) whenever possible.

This library can decode at least 4 billions of compressed integers per second on most desktop or laptop processors. That is, it can decompress data at a rate of 15 GB/s. This is significantly faster than generic codecs like gzip, LZO, Snappy or LZ4.

https://github.com/lemire/FastPFor

https://github.com/lemire/FastPFor/archive/refs/tags/v0.1.8.zip

https://github.com/lemire/FastPFor/archive/refs/tags/v0.1.8.tar.gz

Java May have a use in JS ôo
https://github.com/lemire/JavaFastPFOR

https://github.com/lemire/JavaFastPFOR/blob/master/benchmarkresults/benchmarkresults_icore7_10may2013.txt

*****

SIMDCompressionAndIntersection


C/C++ library for fast compression and intersection of lists of sorted integers using SIMD instructions : https://github.com/lemire/SIMDCompressionAndIntersection

SIMDCompressionAndIntersection
Build Status Code Quality: Cpp

As the name suggests, this is a C/C++ library for fast compression and intersection of lists of sorted integers using SIMD instructions. The library focuses on innovative techniques and very fast schemes, with particular attention to differential coding. It introduces new SIMD intersections schemes such as SIMD Galloping.

This library can decode at least 4 billions of compressed integers per second on most desktop or laptop processors. That is, it can decompress data at a rate of 15 GB/s. This is significantly faster than generic codecs like gzip, LZO, Snappy or LZ4.

*****LZ77*****

Principally an order & load+Vec https://github.com/jearmoo/parallel-data-compression

https://jearmoo.github.io/parallel-data-compression/


Summary of What We Completed

We have written and optimized the sequential version of the Huffman encoding and decoding algorithms, and tested it. For the parallel CPU version of this, we were debating between SIMD intrinsics and ISPC, and OpenMP.

However, Huffman coding compression and decompression doesn’t seem to have a workload that can appropriately use SIMD. This is because there is no elegant way of dealing with bits instead of bytes in SIMD. Moreover, different bytes compress to a different number of bits (there is no fixed mapping of input vector size to output vector size), which makes byte alignment in SIMD very difficult (for example, the compressed form for a random 4 byte input could range from 2 to 4 bytes). This is much worse for decompression, where resolving bit-level conflicts (where a specific encoding spreads over 2 bytes) is almost impossible and might actually result in the algorithm being slower than the sequential version. Therefore, we decided to focus on OpenMP.

For compression, we first sort the array in parallel, to minimize number of concurrent updates to the shared frequency dictionary, reducing contention and false sharing. We also use fine-grained locking for the frequency dictionary, individually locking each key-value pair. Once the symbol codes have been determined, each symbol is replaced by its code, and all symbols are so processed in parallel.

Decompression is inherently sequential, and hence much harder to parallelize. In this case, we take advantage of the self-synchronizing property of Huffman coding, which allows us to start at an arbitrary point in the encoded bits, and assume that at some point, the offset in bits will correct itself, resulting in the correct output thereafter.

We read about the LZ77 algorithm and explored the different variants of the algorithm. We also explored different ways to parallelize LZ77. One naive approach is running the LZ77 algorithm along different segments of the data. This approach could output the same result as the sequential implementation if we use a fixed size sliding window and reread over some of the data. Another approach is the one outlined in Practical Parallel Lempel-Ziv Factorization which uses an unbounded sliding window and employs the use of prefix sums and segment trees to calculate the Lempel-Ziv factorization in parallel.

Update on Deliverables

Our sequential implementations are close to finished, and we have some idea of how to parallelize the algorithms. Our goal for the checkpoint was to have both of these parts finished, but we have not completely met the goal. We may pivot and work on parallelizing the compression and decompression of the Huffman coding algorithm and drop the LZ77 part of the project altogether.

Our new goals:

Parallelize the Huffman Coding compression.
Parallelize the Huffman Coding decompression or LZ77 compression

Hope to achieve:
Both parts of part 2 in our new goals.

*****ZPNG


Huffman source, Requires analysis https://github.com/catid/Zpng

Small experimental lossless photographic image compression library with a C API and command-line interface.

It's much faster than PNG and compresses better for photographic images. This compressor often takes less than 6% of the time of a PNG compressor and produces a file that is 66% of the size. It was written in just 500 lines of C code thanks to Facebook's Zstd library.

The goal was to see if I could create a better lossless compressor than PNG in just one evening (a few hours) using Zstd and some past experience writing my GCIF library. Zstd is magical.

I'm not expecting anyone else to use this, but feel free if you need some fast compression in just a few hundred lines of C code.

**************************

Main interpolation references:


Interpolation https://drive.google.com/file/d/1dn0mdYIHsbMsBaqVRIfFkZXJ4xcW_MOA/view?usp=sharing

ICC & FRC https://drive.google.com/file/d/1vKZ5Vvuyaty5XiDQvc6LeSq6n1O3xsDl/view?usp=sharing

FRC Calibration >

FRC_FCPrP(tm):RS (Reference)

https://drive.google.com/file/d/1hEU6D2nv03r3O_C-ZKR_kv6NBxcg1ddR/view?usp=sharing

FRC & AA & Super Sampling (Reference)

https://drive.google.com/file/d/1AMR0-ftMQIIC2ONnPc_gTLN31zy-YX4d/view?usp=sharing

Audio 3D Calibration

https://drive.google.com/file/d/1-wz4VFZGP5Z-1lG0bEe1G2MRTXYIecNh/view?usp=sharing

2: We use a reference pallet to get the best out of our LED; Such a reference pallet is:

Rec709 Profile in effect : use today! https://is.gd/ColourGrading

Rec709 <> Rec2020 ICC 4 Million Reference Colour Profile : https://drive.google.com/file/d/1sqTm9zuY89sp14Q36sTS2hySll40DilB/view?usp=sharing

For Broadcasting, TV, Monitor & Camera https://is.gd/ICC_Rec2020_709

ICC Colour Profiles for compatibility: https://drive.google.com/file/d/1sqTm9zuY89sp14Q36sTS2hySll40DilB/view?usp=sharing

https://is.gd/BTSource

Colour Profile Professionally

https://displayhdr.org/guide/
https://www.microsoft.com/store/apps/9NN1GPN70NF3

*Files*

This one will suite Dedicated ARM Machine in body armour 'mental state' ARM Router & TV https://drive.google.com/file/d/102pycYOFpkD1Vqj_N910vennxxIzFh_f/view?usp=sharing

Android & Linux ARM Processor configurations; routers & TV's upgrade files, Update & improve
https://drive.google.com/file/d/1JV7PaTPUmikzqgMIfNRXr4UkF2X9iZoq/

Providence: https://www.virustotal.com/gui/file/0c999ccda99be1c9535ad72c38dc1947d014966e699d7a259c67f4df56ec4b92/
https://www.virustotal.com/gui/file/ff97d7da6a89d39f7c6c3711e0271f282127c75174977439a33d44a03d4d6c8e/

Python Deep Learning: configurations

AndroLinuxML : https://drive.google.com/file/d/1N92h-nHnzO5Vfq1rcJhkF952aZ1PPZGB/view?usp=sharing

Linux : https://drive.google.com/file/d/1u64mj6vqWwq3hLfgt0rHis1Bvdx_o3vL/view?usp=sharing

Windows : https://drive.google.com/file/d/1dVJHPx9kdXxCg5272fPvnpgY8UtIq57p/view?usp=sharing

Friday, June 10, 2022

JIT Compiler

Driver & Firmware Integrated JIT Compiler (c)RS

Driver & Firmware Integrated JIT Compiler - DPIC Display Protocol Indirect Compute 2022

Presenting JIT for hardware interoperability & function : https://is.gd/DisplaySourceCode

Integrated JIT Compiler directly into a Shader & OpenCL / Direct Compute Driver Ethernet Protocol Socket & IP

To & from all devices though Firmware Central JIT Compute Compiler

Computation tasks can be carried out by all installed Hardware & USB / Plugged devices:

WebGPU
Python
JavaScript
WebCL, OpenCL & Direct Compute
JIT compiled maths

Indirect Computation such as maths in Application : WebGPU, WebCL, OpenCL & Direct Compute.

Utilising Computation is as simple as having a V8 WebGPU function available,

May be directly available from the GPU without accessing the CPU if SDK is directly supported in GPU RAM...

So in the case of a TV BlueRay Player as an example; We may infact simply be able to integrate..

HTTPS: WebGPU, WebCL, OpenCL & Direct Compute & Methods such as JIT compiled maths.

The plan we use is to; Integrate JIT Compiler directly into a Shader & OpenCL / Direct Compute Driver Ethernet Protocol Socket & IP

Computation tasks can be carried out by all installed Hardware & USB / Pluged devices,

To & from all devices though Firmware Central JIT Compute Compiler

*

Kernel Method requires around 20Kb + Cache Kernel run on OpenCL &or Direct Compute,
Closest device runtime &or Operation infrastructure procedure call.

In the case examples:

Camera focus OpenCL Kernel Ofload
(Edge detect, No image : edges & 4pixels with gradient with jpg compression)

Audio device with buffers OpenCL Kernel Ofload
(processing input is from CPU to Audio Device : Simple Objective Pre Processing case)

SSD & HDD Firmware OpenCL Kernel Ofload
(Location & Write & Math proof of safe write &or read, Error correction)

Printer OpenCL Kernel Ofload
In the case of the printer the postscript driver "Is NOT" installed in your router,
The router prints but has basic drivers,

OpenCL Kernel Ofload (from printer),
Makes the task of processing a Postscript Font & Curl Angle print; Easy!

If you have a USB Hub with processor,
The Postscript Instruction Set is processed as OpenCL Vector Print

*

DPIC Device Protocol Indirect Compute Hub

Proposed HDMI/DisplayPort Hub (also GPU Processed)
Proposed USB Hub,
Proposed Bluetooth Hub,
Proposed WiFi Hub
Proposed Ethernet/Net Hub

with
50Mhz to 800Mhz processor with Dynamic Eco settings
*

On the aspect of HDMI & DisplayPort HTTP Ethernet protocol - DPIC Display Protocol Indirect Compute 2022 (c)RS https://bit.ly/VESA_BT

On the aspect of HDMI & DisplayPort HTTP Ethernet protocol; Several forms of Computation exist as possible for the equipment involved : Televisions, Monitors & GPU & CPU

*

(c)Rupert S https://bit.ly/VESA_BT

Research topic RS : https://is.gd/Dot5CodecGPU https://is.gd/CodecDolby https://is.gd/CodecHDR_WCG https://is.gd/HPDigitalWavelet https://is.gd/DisplaySourceCode

*

Example : JIT Optimise Dynamic code - DPIC Device Protocol Indirect Compute

Audio/Video/GPU/CPU/Urt/USB/BT : hardware to slow or fast? trade Processor Resources : How? DCP:JIT

Camera Focusing API for Web : Application,
Because Computers surely focus a camera better if we use DPIC : Device Compute
Processing JIT Compiler,
Then Latency is not the issue!

Video & Audio can do with additional processing : How? DCP:JIT

Monitor would be able to do so much more! With additional processing : How? DCP:JIT

Kernel Method requires around 20Kb + Cache Kernel run on OpenCL &or Direct Compute,
Closest device runtime &or Operation infrastructure procedure call.

Tier processing; Objectives:

High quality process,
Performance,
Shared workload,
Appropriate Computing unit

In the case examples:

Camera focus OpenCL Kernel Ofload
(Edge detect, No image : edges & 4pixels with gradient with jpg compression)

Audio device with buffers OpenCL Kernel Offload
(processing input is from CPU to Audio Device : Simple Objective Pre Processing case)

SSD & HDD Firmware OpenCL Kernel Offload
(Location & Write & Math proof of safe write &or read, Error correction)

Printer OpenCL Kernel Offload
In the case of the printer the postscript driver "Is NOT" installed in your router,
The router prints but has basic drivers,

OpenCL Kernel Offload (from printer) > (from USBHub : Some) > (Router back to printer),
In an ideal situation the Kernel processes the next tier up; In this case Pro-USBHub;
Leaving the router process free but with a very high quality printing job done.

Makes the task of processing a Postscript Font & Curl Angle print; Easy!

If you have a USB Hub with processor,
The Postscript Instruction Set is processed as OpenCL Vector Print

*
DPIC Device Protocol Indirect Compute Hub

Proposed HDMI/DisplayPort Hub (also GPU Processed)
Proposed USB Hub,
Proposed Bluetooth Hub,
Proposed Wifi Hub
Proposed Ethernet/Net Hub

with
50Mhz to 800Mhz processor with Dynamic Eco settings
*

Inter-device JIT Compiler Kernels (c)RS


JIT Compiler : Driver facing the Monitor is included with JIT Compiler Firmware
subjectively...

For the JIT Compiler to be available add the JIT Compiler to the HDMI
& Displayport Driver,

Facing from the Monitor/AUDIO/VIDEO/BUSS/URT <>
GPU <> CPU

USB & Bluetooth require both the USB, Dongle adapter <> BUSS/URT <> CPU/GPU...
To further connect Printers & other devices...

Under the same principle the Lens of a camera operating under the mounting fixture requires a fast connection,
In order to utilize Infrared/UV/Laser & Light or DIODE Controlled fixtures that require special firmware downloaded kernels,

These kernels are flexible & will speed up devices & assure top performance with:

16Bit, 32Bit, 64Bit & Float Kernels

Driving the monitor & Learning sharper graphics

https://is.gd/BTSource


Firstly, we need an OpenCL Kernel : PocCL :

PoCL Source & Code
https://is.gd/LEDSource

MS-OpenCL
https://is.gd/MS_OpenCL

HIP_CUDA on OpenCL & SPIRV
https://github.com/CHIP-SPV/chipStar

HIP_CUDA on OpenCL & SPIRV : ZIP
https://is.gd/HIP_CUDAonOpenCL

X86Features-Emu
https://drive.google.com/file/d/15vXBPLaU9W4ul7lmHZsw1dwVPe3lo-jK/view?usp=usp=sharing

*

Code/JS/OpenCL/Machine Learning Processing Block Size Streamlining (c)RS


Dataset AV1/VP9/MPEG/H265/H264 : case example
My personal observation is that decompression & compression performance relates to block size & cache

SiMD 8xBlock x 8xBlock Cube : 32Bit | x 4 128Bit | x 8 256Bit | x 16 512Bit
Cache Size : 32Kb Code : Code has to be smaller inline than 32Kb! Can loop 4Kb x 14-1 for main code segment

Cache Size 64Kb Data : Read blocks & predicts need to streamline into 64Kb blocks in total,
4Kb Optimized Code Cache
4Kb Predict (across block for L2 Multidirectional)
16Bit Colour Compressed block 4x16Bit (work cache compressed : 54Kb
Lab Colour ICC L2 & block flow L2

https://science.n-helix.com/2022/09/ovccans.html

*

Combining JIT PoCL with SiMD & Vector instruction optimisation we create a standard model of literally frame printed vectors :

VecSR that directly draws a frame to our display's highest floating point math & vector processor instructions; lowering data costs in visual presentation & printing.

(documents) JIT & OpenCL & Codec : https://is.gd/DisplaySourceCode

Include vector today *important* RS https://vesa.org/vesa-display-compression-codecs/

https://science.n-helix.com/2022/06/jit-compiler.html

https://science.n-helix.com/2022/04/vecsr.html

https://science.n-helix.com/2016/04/3d-desktop-virtualization.html

https://science.n-helix.com/2019/06/vulkan-stack.html

https://science.n-helix.com/2019/06/kernel.html

https://science.n-helix.com/2022/03/fsr-focal-length.html

https://science.n-helix.com/2018/01/integer-floats-with-remainder-theory.html

https://science.n-helix.com/2022/08/simd.html

*****

Good stuff for all networks nation wide, the software is certificate signed & verified
When it comes to pure security, We are grateful https://is.gd/SecurityHSM https://is.gd/WebPKI
TLS Optimised https://drive.google.com/file/d/10XL19eGjxdCGj0tK8MULKlgWhHa9_5v9/view?usp=share_link
Ethernet Security https://drive.google.com/file/d/18LNDcRSbqN7ubEzaO0pCsWaJHX68xCxf/view?usp=share_link

These are the addresses directly of some good ones; DNS & NTP & PTP 2600:c05:3010:50:47::1 2604:6600:2000:45::1 2604:6600:2000:2b::1 2a06:98c1:54::c12b 142.202.190.19 172.64.36.1 172.64.36.2 108.181.201.22 108.181.165.159

Friday, April 1, 2022

VecSR - Vector Standard Render

VecSR - Vector Standard Render


VESA Standards : Vector Graphics, Boxes, Ellipses, Curves & Fonts : Consolas & other brilliant fonts : (c)RS

Vector Compression VESA Standard Display protocol 3 : RS

SiMD Render - Vector Graphics, Boxes, Ellipses, Curves & Fonts

*
VecSR (c)RS

VecSR is the principle for SiMD to accomplish a 2D & 3D trace of Rays & Vectors,
Principally the technology can do a couple of things (and more):

Vectorising the Instruction & Presentation functions

Precisely upscale (presentation)
Save data bandwidth on connections for monitors & printers & mice; By 'Vectorising the Instruction'
Present fonts & vectors to infinity..
Present wavelets in their Ultimate Vector form..

There is no limit to Precision & Cache; Because presentation can be 'Dynamic Cache' & Precision is upto the Bit Depth of the hardware presenting.

32Bit, 16Bit SiMD presents 32Bit, 16Bit FP16b, So precision is not a problem; Vectors are presented full precision as we want.
*

*
32Bit SiMD Operations Available on AVX Per Cycle (A Thought on why 32Bit operations are good!)
(8Cores)8*32Bit SiMD(AVX) * 6(times per cycle) * 3600Mhz = 1,382,400 Operations Per Second

Security Relevant Extensions
SVM : Elliptic Curves & Polynomial graphs & function
AES : Advanced Encryption Standard Functions
AVX : 32Bit to 256Bit parallel Vector Mathematics
FPU : IEEE Float Maths
F16b : 16Bit to 32Bit Standards Floats
RDTSCP : Very high precision time & stamp

Processor features: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 htt pni ssse3 fma cx16 sse4_1 sse4_2 popcnt aes f16c syscall nx lm avx svm sse4a osvw ibs xop skinit wdt lwp fma4 tce tbm topx page1gb rdtscp bmi1

a
 \__c
  |
  b
Now Matrix Vectors for vector rendering https://science.n-helix.com/2023/06/map.html


Photos & Performance https://is.gd/4447GamerWEBB



VecSR Anticipated Quadratic Array : Font Rendering
https://gpuopen.com/learn/mesh_shaders/mesh_shaders-font_and_vector_art_rendering_with_mesh_shaders/

*

OT-SVG Fonts & TT-SVG Obviously Rendered in Direct X 9+ & OpenGL 3+ Mode & Desktop Rendering modes


Improve Console & TV & BIOS & General Animated Render

Vector Compression VESA Standard Display protocol 3 : RS

SiMD Render - Vector Graphics, Boxes, Ellipses, Curves & Fonts
Improve Console & TV & BIOS & General Animated Render

Vector Display Standards with low relative CPU Weight
SiMD Polygon Font Method Render

Default option point scaling (the space) : Metadata Vector Fonts with Curl mathematical vector :

16 Bit : SiMD 1 width
32 Bit : SiMD Double Width

High precision for AVX 32Bit to 256Bit width precision.

Vectoring with SiMD allows traditional CPU mastered VESA Emulation desktops & safe mode to be super fast & displays to conform to VESA render standards with little effort & a 1MB Table ROM.

Though the VESA & HDMI & DisplayPort standards Facilitates direct low bandwidth transport of and transformation of 3D & 2D graphics & fonts into directly Rendered Super High Fidelity SiMD & AVX Rendering Vector

Display Standards Vector Render : DSVR-SiMD Can and will be directly rendered to a Surface for visual element : SfVE-Vec

As such transport of Vectors & transformation onto display (Monitor, 3D Unit, Render, TV, & Though HDMI, PCI Port & DP & RAM)

Directly resolve The total graphics pipeline into high quality output or input & allow communication of almost infinite Floating point values for all rendered 3D & 2D Elements on a given surface (RAM Render Page or Surface)

In high precision that is almost unbeatable & yet consumes many levels less RAM & Transport Protocol bandwidth,

Furthermore can also render Vector 3D & 2D Audio & other elements though Vector 'Fonting' Systems, Examples exist : 3D Wave Tables, Harmonic reproduction units for example Yamaha and Casio keyboards.

RGBA Composite Layer X-OR


RGBA Can simply be the shape printed onto alpha layer; Wide Transparency effect.
RGB-Supposition is X-OR Shape on mapping block or cube or curve & shape; Due to Alpha Alias smooth blending is achieved.

*

Furthermore can also render Vector 3D & 2D Audio & other elements though Vector 'Fonting' Systems, Examples exist : 3D Wave Tables, Harmonic reproduction units for example Yamaha and Casio keyboards.

Personally QFT is a much more pleasurable experience than VRR at 2xFPS+
Stable FPS & X-OR Partial Frame Retention saving on compression.

"QFT a Zero compression or low level compression version of DSC
1.2bc

X-OR Frame Buffer Compression & Blank Space Compression:
Vector Compression VESA Standard Display protocol 3"

"QFT transports each frame at a higher rate to decrease “display
latency”, which is the amount of time between a frame being ready for
transport in the GPU and that frame being completely displayed. This
latency is the sum of the transport time through the source’s output
circuits, the transport time across the interface, the processing of
the video data in the display, and the painting of the screen with the
new data. This overall latency affects the responsiveness of games:
how long it appears between a button is pressed to the time at which
the resultant action is observed on the screen.

While there are a lot of variables in this equation, not many are
adjustable from an HDMI specification perspective. QFT operates on the
transport portion of this equation by reducing the time it takes to
send only the active video across the cable. This results in reduced
display latency and increased responsiveness."

*

(c)Rupert S

(QT_SECC) ECC Temporal Tick for low energy devices & computer systems : RS


(including GPU & RAM & Fast Storage),
Fast & high performance Elliptic Curves 8Bit to 128Bit

Ideal standards of 16Bit Elliptic curves for Audio, Video, 3D Texture & Edge shaping...
As described here we create edges & cubes & fills & Obviously Elliptic Curves!

We can shape digital audio directly; But also Video & Textures; Any shape that matches our description..
Any dream involving a precisely defined maths object that is a shape vector.

This is not just a security device.

RS

Direct Rendering Matrix Vectors (c)RS


a
 \__c
  |
  b
Now Matrix Vectors for vector rendering https://science.n-helix.com/2023/06/map.html

DRMV Direct Surface Draw : Laser Printers, Screen, GPU, CPU & Applications of DirectX, Vulkan & OpenCL & Direct Compute HTML5 & JS Buffer

With SiMD & Neon & AVX Features common to CPU & GPU, We can directly compose Vectors & Texture compose directly to the screen..

Using Matrix Formula Maths : a, b, c, 3D render; We are not simply limited to enhanced eliptic curve & cubic functions..

Optimised Eliptoid, Elliptic & Eccentric cuboid functions significantly improve a VESA Certified Render,

QNON, Ellipto-centric force physics & dimensional realities; Become a Vector Render Reality Matrix :

Holograms & Vector Drawing with SiMD, AVX, Matrix Units & FPU or Integers with RollINT.

We can draw Squares, Cubes, Curves, Ovoids, Ellipsoids, Shapes & Voxels & Tixels directly to any renderable surface; Including VESA Approved Monitor standards to VecSR Standards..

Directly from any available FPU, SiMD, Float or Integer unit.. Directly to any Video & Audio Buffer,
Therefore directly to Vector drawing surfaces such as: Laser Printers, Screen, GPU, CPU & Applications of DirectX, Vulkan & OpenCL & Direct Compute HTML5 & JS Buffer

For reference to the functions of Curves, Elliptic & Cuboids that DRMV can run:
Architecture Fast Instructions for FMA https://science.n-helix.com/2023/06/map.html

Rupert S

Reference operators


Vector Font Render Sources{

https://github.com/GreenLightning/gpu-font-rendering

https://github.com/azsn/gllabel

https://github.com/KeinR/Etermal/blob/master/README.md

};



Meshlet compression & our example : RS


https://gpuopen.com/learn/mesh_shaders/mesh_shaders-meshlet_compression/

Meshlet compression & our example : RS

Now as explained in the document; You subtract one edge define by reusing one edge,

Defining a square with squares around it; With 4 squares; You can define the central square & define the other squares to be of the same size,

Squares thus defined:
________
|__|__|__|
|__|__|__|
|__|__|__|

As you can see the square only needs to define the original square (in a polygon sense),
You can copy the source square from cache!

You can copy triangles from the cache also! If that polygone is a maths shape such as square; You can define a group! & cache that!

If that shape is a fractal? You can define the origin (regular fractals do that,..

However; Some fractals are fast! What fractals are those?

Fake balls : Fractal curves,
Shapes such as ellipses that have an equal size and symmetry (one side, 2 sides, All sides)
Evenly distributed curves, As seen from the center...

Common shapes are : square, Line, cube, oblong, oval, circle, letters & numbers & such content as hash # & symmetrical patterns..

"SiMD & AVX instanced array"

Hardware Maths Accelerated instanced_arrays 03/10/2024 1337 https://is.gd/DictionarySortJS

https://science.n-helix.com/2022/04/vecsr.html

(c)Rupert S

*

BT-2.4G QT_SECC in context of VECSR


Able to be used for Motion, Haptic, Video, Texture, Audio wavelet creation & use:

The Wave pattern principle is in principle a content of pure colour curves, both depth & content of pixel,

But also a means by which elliptic curves are created with great simplicity..
So that singular hardware like F16 SiMD can truly create a master piece; Both Crypto & Dimensional 'art'

BT-2.4G Quartz Time Crystal Tick Simple Elliptic Curve to Support FIPS 128Bit on Unifier USB,

Modulation to 16Bit& 32Bit & 64Bit & 128Bit allow for different types of SiMD & AVX
Allow for Android & Linux & Windows; ARM & X86 & GPU Processors

Presented with a single tick \_/-\_/ Complex modulating Elliptic curves of 8Bit & 16Bit & 32Bit & 64Bit & 128Bit lengths,

16Bit to 64Bit & 128Bit output curves; Through temporary ECC certificate..; Additionally ChaCha_Poly & AES Ciphers..

Rupert S

Principle:
Bluetooth dongle LE Protocol https://drive.google.com/file/d/17csRnAfdceZiTSnQZvhaLqLSwL__zsIG/view?usp=sharing

https://science.n-helix.com/2023/06/map.html
https://science.n-helix.com/2023/02/smart-compression.html
https://science.n-helix.com/2022/04/vecsr.html

https://science.n-helix.com/2022/03/ice-ssrtp.html
https://science.n-helix.com/2021/11/ihmtes.html

https://science.n-helix.com/2022/08/jit-dongle.html
https://science.n-helix.com/2022/06/jit-compiler.html

*

How JS & WebASM use a,b,i,c maths improves the total speed of application load & mouse control


For reference to the functions of Curves, Elliptic & Cuboids that DRMV can run:
Architecture Fast Instructions for FMA

https://science.n-helix.com/2023/06/map.html

https://science.n-helix.com/2022/04/vecsr.html

FMA AVX Performance table: 2Flops per Cycle per FMA Unit
Architecture Fast Instructions for FMA

Reference Tables https://www.uio.no/studier/emner/matnat/ifi/IN3200/v19/teaching-material/avx512.pdf

Operators in C
● Arithmetic
a + b, a – b, a*b, a/b, a%b
● Bitwise
a | b, a & b, a ^ b, ~a
● Bit shift
a << b, a >> b (signed), a >> b (unsigned)
● Logical operators
a && b, a || b, !a
● Comparison operators
a == b, a != b, a < b, a <= b, a > b, a >= b
● Tertiary operator
x = a ? b : c
● Special functions:
sqrt(x), abs(x), fma(a,b,c), ceil(x), floor(x)

Fast division for constant divisors

Calculate r = a/b where b is a constant
With floating point we precompute (at compile time
or outside of the main loop) the inverse ib = 1.0/b.
r = ib*a
Floating point division with constant divisors
becomes multiplication
With integers the inverse is more complicated
ib,n = get_magic_numbers(b);
r = ib*a >> n

Integer division with constant divisors becomes
multiplication and a bit-shift

Fast Division Examples
● x/3 = x*1431655766/2^32
27*1431655766/2^32 = 3
● x/1000 = x*274877907/2^38
10000*274877907/2^32 = 10
● x/314159 = x*895963435/2
7*314159*895963435/2^48 = 7

Dividing integers by a power of two can be done with a bit shift which is very fast.

RS

High speed Per operation Cycle operations of D R² Pi

An (A[diameter]*B²[Pi] : D * R² operation is 2 Cycles, this specialised Arc, Sin, Tan operation can be accomplished a couple of ways in a single cycle,

Options table : D R² Pi

Firstly by sideways memory load in lower Single Precision to double precision output in a SiMD

You need to pre cache R²You can use the same value for R or for D &or both
You can pre cache all static D &or R, So you can vary either D or R & single cycle
You need to perform 2 operations , Diameter & R² & obviously they are relational!

For examples:

R = Atom Zink (standard size!) Cache D R
You move a compass but the needle is the same size! Cache D
You draw faces but the width is the same, Cache D
You draw faces but the Shape is the same but size is not! Cache R

Rupert S

https://en.wikipedia.org/wiki/FMA_instruction_set
https://en.wikipedia.org/wiki/Advanced_Vector_Extensions
https://en.wikipedia.org/wiki/AArch64#Scalable_Vector_Extension_(SVE)

High-Performance Elliptic Curve Cryptography: A SIMD Approach to Modern Curves
https://www.lasca.ic.unicamp.br/media/publications/FazHernandez_Armando_D.pdf
https://science.n-helix.com/2023/06/map.html
https://science.n-helix.com/2022/04/vecsr.html

Updated JS Sourcery to be found at https://is.gd/LEDSource

(Simple Install) Website Cache JS Updated 2021-11 (c)RS https://bit.ly/CacheJS
(Simple Install) Science & Research Node High Performance Computing Linux & Android https://is.gd/LinuxHPCNode

Presenting JIT for hardware interoperability & function : https://is.gd/DisplaySourceCode


(Simple Install) Website Server Cache JS Updated 2021-11 (c)RS https://bit.ly/CacheJSm
(Simple Install) Website Server Cache JS Work Files Zip Updated 2021-11 (c)RS https://bit.ly/AppCacheJSZip

https://npm.n-helix.com/bundles/

Python Deep Learning:

AndroLinuxML : https://drive.google.com/file/d/1dVJHPx9kdXxCg5272fPvnpgY8UtIq57p/view?usp=sharing
Linux : https://drive.google.com/file/d/1u64mj6vqWwq3hLfgt0rHis1Bvdx_o3vL/view?usp=sharing
Windows : https://drive.google.com/file/d/1dVJHPx9kdXxCg5272fPvnpgY8UtIq57p/view?usp=sharing

Andro-linux libs : x86 & ARM : Learn
https://drive.google.com/drive/folders/1BRQOIK1eAUEMnTTGjsQ0h0g6jGLzWqZI

good stuff for all networks nation wide, the software is certificate signed & verified
When it comes to pure security, We are grateful https://is.gd/SecurityHSM https://is.gd/WebPKI
TLS Optimised https://drive.google.com/file/d/10XL19eGjxdCGj0tK8MULKlgWhHa9_5v9/view?usp=share_link
Ethernet Security https://drive.google.com/file/d/18LNDcRSbqN7ubEzaO0pCsWaJHX68xCxf/view?usp=share_link

These are the addresses directly of some good ones; DNS & NTP & PTP 2600:c05:3010:50:47::1 2607:fca8:b000:1::3 2607:fca8:b000:1::4 2a06:98c1:54::c12b 142.202.190.19 172.64.36.1 172.64.36.2 38.17.55.196 38.17.55.111

*

Drawing tools & functions that are the basis of our draw frame & font functions : Polygon maths


Core Processor features : SVM, SiMD, FPU
Core tools : https://science.n-helix.com/2019/06/vulkan-stack.html

Reference material for Drawing Elliptoids, Curves & Polygons

SVM Elliptic Curve magic:
Fractal maths for improved efficiency & Combustion energy, Regard the photos & the FX8320E for details

Effective Application of SVM Processor Elliptic Maths
https://is.gd/SVMefficiency

Linear Bounding Volume Hierarchy &
Elliptic Bounding Volume Hierarchy for SVM Processor Feature:
SVM Can be emulated in SiMD pure 32Bit Single or 64Bit Double Precision,
& is for high complexity rendering such as non regular windows.

https://www.phoronix.com/scan.php?page=news_item&px=RADV-LBVH-Lands

SVM Can be emulated in SiMD pure 32Bit Single or 64Bit Double Precision..
Is useful for creating non Circle curves such as elliptoids & oblong wave boxes.

In VSR & VSR Variable Lighting we can define spaces with eliptoids SVM,
Therefore shape around trees & grasses & animals &or people & Whales.

https://www.youtube.com/watch?v=UojqzrPtR70

(c)RS

*

FFT or QFFT : Fast Fourier Transform


FFT or QFFT is not only about audio; But also Video & 3D, Mouse & input/output devices (c)RS 2022

FFT or QFFT is not only about audio; But also Video & 3D,
In fact FFT Fast Fourier Transforms are about any device such as a mouse that directly interacts with Waves,

Such a device is the laser mouse & pointer; The primary reason is to use Noise reduction & path smoothing,
Primarily to create a 16Bit to 256Bit pure float with high compression or pack bit properties.

Creating Sine-oidial curves & waves or SiMD, Float & packed integer/Float operations saves on bandwidth & increases messaging speed therefore!

Both the input & output from Bluetooth, 2.4G & USB & Serial can in fact be reduced to mapped Curves & angles; While this introduces a small error factor & this is a factor that producers & driver developers need to work out & create error margins for.

Creation & development of Ultra high precision Input & output for Humans, Robots & precision pointers; Requires a precise production FFT & to account for the fact surrounding the interactive motion of point A to point B; & In fact point C...

Development continues & today's mission is to open minds about why we use FFT & noise reduction & Curve maps such as elliptic SVM & Bit Averaging Fast transforms for Center point Algebra & Math Tables & Graphs.

Further study includes Raytracing & All Haptic motion; Sensors & Car engine Mechanics.

(c)Rupert S

*

Include vector today *important* RS https://vesa.org/vesa-display-compression-codecs/

https://science.n-helix.com/2016/04/3d-desktop-virtualization.html

https://science.n-helix.com/2019/06/vulkan-stack.html

https://science.n-helix.com/2019/06/kernel.html

https://science.n-helix.com/2022/03/fsr-focal-length.html

https://science.n-helix.com/2018/01/integer-floats-with-remainder-theory.html

https://bit.ly/VESA_BT

https://science.n-helix.com/2023/03/path-trace.html

https://science.n-helix.com/2023/02/smart-compression.html

https://science.n-helix.com/2022/09/ovccans.html

https://science.n-helix.com/2022/08/simd.html

*

Core Concepts of Direct Vector Render Frame Buffers & Cache


LHP_DSC_Xor : Screen Fast Buffer Access


VESA Standard Ethernet Standard Frame Protocol for QFT, VRR & Low Latency High Performance Dynamic Compression XOR Frame Refresh : LLHP_DSCX : LHP_DSC_Xor

QFT & VRR basically allow the TV to float a resolution refresh free from Frame Cache Memory Refresh (Refueling the Cache Buffer) ,
Basically the frame can be fetched from the Frame Cache (4MB to 64MB) Without interacting with the CPU

This means a Fast Direct DMA Cache pull on frame to Screen & does not demand that the CPU need to perform this fast; Additionally the Frame comes without tearing or Frame pulls from the HDMI or display port VESA Ethernet Standard Frame Protocol.

Rupert S

*

3D DR_LC : 3D Layers to Direct Render Layer Composing : OS, DSC, Codecs, DirectX & Vulkan


https://science.n-helix.com/2022/10/ml.html

https://science.n-helix.com/2022/04/vecsr.html

Here are the Operation Processor Extensions available to EdgeTPU:
https://coral.ai/docs/edgetpu/models-intro/

The sample examples show what a powerful specialised instruction set can do!
To explain more; The EdgeTPU is a Matrix multiplier & Adder..

The instructions such as transpose allow for example mapping one image on another for difference detection...
Flexible uses for each function can literally be based on the basic concept of the instruction,

Basic assumptions lead to convoluted & complex examples..

Examples that are required to do such things as check one bitmap for identical content (in effect XOR)

Quantize : image pixels..

Max Min Mean : Dithering or gaussian blends (complex XOR & layering or edge feathering) & more!

StridedSlice & Slice : partition a frame into parts to render in a grid; slice CSS isolated content in rendering.

SpaceToDepth : Dynamically allocate depth layers to single frame content such as text boxes or photos..
So why ? so we can Dither edges & Fonts & minimise ram usage to dynamic content.

We AveragePool2d to gaussian blend the layers together; In principle we average weight the layers to a final,
Single layer; DSC VESA

Alternative is to Paint major content on a single layer involving the CSS backdrop..
Moving content on top of it on a secondary layer; Makes sense to me! speed wise,

We could use an average pool with Weights (+10/30/100 to -10/30/100) & Feather & Gaussian blend down if we like!

MaxPool2d define layer amount.

ResizeBilinear, ResizeNearestNeighbor : Resize textures for appropriate size of mouse pointers & cursors & content pictures or video..

DepthwiseConv2d : we can down convert layers to 2D, in principle in chrome we convert layered textures during final frame generation to a single flattened layer..

Transpose : layers folded into a single frame render fast! bear in mind that we have to HOLD THE LAYERS in a single fetch!
Buffer optimization to ram size required.. Memory optimization is crucial to hold all layers in a single fetch.

Alternatively combine layers with Transpose & DepthwiseConv2d combined.

In a genuine way layering mouse pointers ontop of DSC frames makes a lot of sense in the terms of response & compression..

you have to think in terms of knowing what is under that mouse pointer; there are two frames of reference to this:

deliberated previous frame forward predict with icon buffer to load over it (A small texture)

Layered responses; in layered responses the Processor processes a java script css layer in the form of the operating system & vectors..

The motion pointer or animation travels over the top in a secondary layered response,
The formation of layers lowers processing costs & speeds up UI response timers; lowering overall compression costs because the first layer is fully converted into an almost static prediction response &or reactionary differentiator system..

The mouse pointer floats over a texture (a Desktop CSS for example, white box for example), The DSC codec encodes pointers (vectors) specific to the first & second layer..

we increase the accuracy of prediction vectors by knowing the desktop layer first & knowing if it is animated or static; we also prefetch the animation & time sync it correctly; so that it animates properly..

We seed the mouse motion vectors & animate the pointer vectors over the top of the Desktop layers..

Desktop composure list

1 Mouse pointer
2 Icons in box
3 Desktop

Vector priority list:

1 desktop
2 Icons & frames
3 mouse pointer.

(c)RS

Packed Bit Z-Buffer


The truth is that simple features & simple layered maths make a small Z-Buffer a logical choice,

Bear in mind that a small Z-buffer can be done between 4Bit & 32bit & is readily handled by the 4Bit or 8bit extensions; Such as RISC V & ARM..

the logical choice being Packed bit u32/8 for example; where multiple layers can be handled in a single fetch & present cycle.

It makes sense to offer a Z-buffer render for Windows, Linux, Android, Consoles, 
For desktop & UI rendering in particular : https://science.n-helix.com/2022/04/vecsr.html

RS

QFT Quick Frame Transport : Motion Vectors


The 3D Layers to 2D layers & flattening single layer (DSC Compression or codec) is a good system.. with lower latency,

Send Vector Predictions from the rendering frame : Statistical Leveraged Future Motion Vector Prediction

Ideally Prediction vectors are produced from at least 2 layers; for example mouse pointer animations L2 & Desktop L1...

This is because if the desktop is not in motion then the vectors are predicting a static content! The mouse however is predicted by being in motion in a clean fashion (in the previous frame),

Because mouse motion is an example where predicting motion is not always easy..
We know ML would define an objective for a prediction such as to an application...

Cross hairs are the same; identified target of motion ML...

So how do we predict motion ? The renderer is preparing the next frame; the pointer or cross hair is in motion in the frame being made!

Most likely we shall be able to send Vector Predictions from the rendering frame,
Layers such as desktop & file icons are in motion or not in the frame being prepared!

Prediction vectors are thus rendered ahead & QFT Quick Frame Transport allows us to send Prediction Vectors early in the frame creation...

Such a thing is called a Statistical Leveraged Future Motion Vector Prediction & is statistical or created in advance during frame production,

QFT Quick Frame Transport is leveraged with motion vectors.

(c)RS

FreeSync & Advanced Sync : GTG 0.1- : QFT : Quick Frame transport


I think implementing QFT & Installing DSC Codec into modes such as 2x Frame transport would work,

Because there would be 2x as much bandwidth for drawing such things as mouse pointers & crosshairs or screen painting content : PSPC :

When you do (RT) Real Time frame transport QFT Quick Frame Transport & VRR will be distributing multiple frames compressed; That way the Display Prediction Vectors update the display very quickly at a much lower bandwidth...

The content strategy is layered with motion vectors for each sub classification in layers.

(c)RS

*

Predicted Content Compression Frame Negotiation (c)RS


Compression for HDMI & DP : VRR & QFT with frame content prediction & Minimal Adjust; X-OR Content replacement

Compression Implicitly supported : STC, DXT, EAC & ATSC & DSC , Most of these compression forms are available in ARM, AMD, NVidia & Intel Hardware & therefore directly supported by us in creating the best frames & video; HDR WCG RGBA/X 4 Channel.

Compression required for a display; Common details include using Compression as a last desperate measure to improve bandwidth for displays on High Definitions such as 4K on HDMI 2!

My personal strategy is to implement compression that is transparent; Starting right at almost non,

Frequently the problem with VRR & QFT is that a frame is sent or not sent...

By utilizing Prediction in compression we force the prediction of an exact copy of present data,
We adjust the frame with X-OR & modify only a few details; Therefore we do not need to send a lot of data & can send more frames!

*

HDMI Input compression : Checker Board 2 frame compression with LZ Compression styles


The application of GZIP Brotli ZSTD compression to screen data tunnels, Allows for 11K for connections on DisplayPort & HDMI,
With the simple switch to automatically lossless compression tunnels,

The use of Checker Board 2 frame compression with LZ Compression styles allows most generic CPU to Deinterlace Double Scan data layers..

Doubling effective resolutions.

QFT Quick Frame Transport in relation to HDMI Input compression:


When you transmit serial frames with the same data compression comes in handy!
So enabling Brotli/ZSTD/GZip/DSC compression with Proofs of frame exact copy or slight modifications..

Now transmit each part of the frame that is exactly the same as a compression copy,

So in effect the frame is micro copied & each part is identified as part of the main frame repeat or new,

In addition if the colour shifts but not the edges or shape; Most of the compression works in reference to HDMI Input compression,

Brotli/ZSTD/GZip/DSC compression works fine in referencing colour shifting light or shape shifting but same light,

Compression works fine.

QFT with SSRTP is perfect for Web+ content refreshing 'Audio & Video' HDMI & VESA DisplayPort connection configurations.

Aligned Byte Codes with 16bit compression codes ZSTD saves 80% of all data costs to content,
Small Byte dictionary compression saves 80% of transmit bandwidth.

(c)RS

Bluetooth dongle LE Protocol https://drive.google.com/file/d/17csRnAfdceZiTSnQZvhaLqLSwL__zsIG/view?usp=sharing

https://science.n-helix.com/2022/03/ice-ssrtp.html

*

The point of Brotli-G is that it minimises the network capacity needed for firmware updates or internet access, most devices use ethernet or wifi; however supporting Brotli-G is going to be fast!

Additionally Brotli-G will allow compressed frames & sub-frames to be compressed flexibly,

What DSC Allows in the form of sub-frames? However Brotli-G Allows sub-Frames...

|FRAME        FRAME|
|SF|SF|SF|SF|SF|SF|SF|

As you can see the intention of sub-framing is to initiate a small section of the screen during the refresh cycles available to QFT & Fast Frame Transport,

We thereby refresh only a small segment & can speed up the process!
Compression is required for efficient sending & we therefore will be using the suggested micro frame format : Brotli-G with Auto Encoding.

Rupert S

VESA + HDMI : Fast Pack Huffmans, Brotli, AutoEncoder https://is.gd/WaveletAutoEncoder https://github.com/GPUOpen-LibrariesAndSDKs/brotli_g_sdk

https://is.gd/CJS_DictionarySort

Python & JS Configurations
https://is.gd/DictionarySortJS

*

Vector Compression VESA Standard Display protocol 3 +

DSC : Zero compression or low level compression version of DSC
1.2bc

Frame by Frame compression with vector prediction.

Personally, QFT is a much more pleasurable experience than VRR at 2xFPS+
Stable FPS & X-OR Partial Frame Retention saving on compression.

X-OR Frame Buffer Compression & Blank Space Compression:

X-OR X=1 New Data & X=0 being not sent,
Therefore Masking the frame buffer,

A Frame buffer needs a cleared aria; A curve or ellipsoid for example,
Draw the ellipsoid; This is the mask & can be in 3 levels:

X-OR : Draw or not Draw Aria : Blitter XOR
AND : Draw 1 Value & The other : Blitter Additive
Variable Value Resistor : Draw 1 Value +- The other : Blitter + or - Modifier

*

PCCFN


The idea Behind PCCFN is to modify the frame by a smaller amount with low bandwidth & thereby increase frame rate by the following method:

DSC Compression is used & Predict is enabled..
Predict is used to redisplay the frame on the screen; With no data needing to be sent : X-OR..
However Modifications are made to the frame by overruling parts of the Static frame with data..

The effect is that only parts of the frame (Vector Motion Prediction); Are sent,

Both bandwidth & speed are preserved & the same effect works from BFrames & Partial Full Frames.

https://hdmi.org/spec21sub/variablerefreshrate
https://hdmi.org/spec21sub/quickframetransport

*

ITS_DHDR_VRR : Gaming & Desktop : HDR, Source-Based Tone Mapping (SBTM)

High Efficiency DSC Screen Dynamic Shift State Screen blanking Replacement
Low Bandwidth Requirement for 40Hz to 240Hz+

HDR, HDMI & Display-port Standards VESA 2022 : Independent Thread
Asymmetric Compute Frame Buffer Tree for HDR, Display & Compression
DSC : RS (c)Rupert S

Composer Frame DSC is where we Compose a frame in the renderer, That
frame is for example the window task bar & another box for the
Explorer frame; The example is not OS Exclusive; Is an example.

We implement DSC Display compression in the frame (smaller than the
display resolution or super sampled),

Every piece of content in the Main Render Frame to HDMI & Display port
is computed independently with static content not being adjusted or
recompressed until needed,

Our goal is to place Every frame or window in a Sub-Buffer Cache & Render to the main Frame Cache/Buffer,

On completion of the frame at whatever FPS Refresh we desire for the Main Frame Buffer,
Effectively we Blitter &or Byte-swap our Window Frame Buffer to a location within the Main Frame buffer,

The location of our window & our localised processing mean that content of each window & therefore process is independently proven to be the Same as the frame before (We X-OR),

Therefore we Frame Predict (DSC) That a small portion of the main frame buffer has the same data,
We do not need to change a thing & so we do not need to utilize the processor to render it..

However if data has changed; Then the change is localised to a single small render space in the main frame buffer & we therefore can refresh the screen faster & Frame Prediction (Like JPG & MPEG)

Proves that we only need to inform the Screen (HDMI & DP Signal in our case);
That no additional date is sent; However any changes to the main frame buffer such as main view or video or text files or HTML Refresh will be Sent & Rendered,
Without Latency issues or large amounts of data being sent though the Cable..

But we still render faster than recompressing a main frame buffer completely & in addition change what we wish per thread without the resulting processing Hanging or waiting on Data To arrive from a baton-pass.

Our reasoning is that each frame is independent; Therefore we compose
in GPU or CPU & independently Compress the Frame within adjusted
context of the HDMI & DisplayPort,

3 Frame Buffer; We can optimise the whole frame with Prediction
Compression if we wish,

The Main goal : Independent Thread Render for Sub-Framing High Dynamic
Range with Independent Application Variable Refresh Rate :
ITS_DHDR_VRR.

The main advantages are : Task bar is Low CPU Resource use but high
refresh rate; low data modification rate over a tiny area of the task
bar,

The Game Window & the Frame (Mostly Square) are drawn with sub-pixel
precision on location..
But the frame that barely changes does not need recompression in DSC..

The Game window does not need to compute or adjust content Compression
for the frame...

Every piece of content in the Main Render Frame to HDMI & Display port
is computed independently with static content not being adjusted or
recompressed until needed.

This works with the HDR, HDMI & Display-port Standards VESA

(c)Rupert S

*

Elliptic Curves & JPEG & MP4/ACC Presentation


Ok so principally we want to create curves with ARC, Sin & Tan,
We can obviously present a curve in 16Bit or even 8Bit; So we can present a curve at the precision we have in the processor (such as 16Bit/32Bit SiMD),

By presenting a curve at higher precision; We can upscale or super sample it,

Super Sampling is principally presenting a curve at higher precision &or softening it with analogue/Digital filters..

So by this example we present a case for elliptic curves presented within the scope of 16Bit or higher SiMD & Floats..

The key idea is that we can use them!

So we can present JPEG, ACC, MP4 as Elliptic curves for upscaling...
We can use Elliptic curves for encryption or presentation on GPU or other processors,
We can present curves to the pixels of a screen surface the same way; scaling them into higher precision.

How well defined that curve is depends on our precision capacity; But we can still use Elliptic curves at any precision we have available.
So what do we want to use Elliptic curves to present ? Anything we need.

RS

*

*Application of SiMD Polygon Font Method Render

*3D Render method with Console input DEMO : RS

3D Display access to correct display of fonts at angles in games & apps without Utilizing 3rd Axis maths on a simple Shape polygon Vector font or shape. (c)Rupert S

3rd dimensional access with vector fonts by a simple method:

Render text to virtual screen layer AKA a fully rendered monochrome, 2 colour or multi colour..

Bitmap/Texture,

Due to latency we have 3 frames ahead to render to bitmap DPT 3 / Dot 5

Can be higher resolution & we can sub sample with closer view priority...

We then rotate the texture on our output polygon & factor size differential.

The maths is simple enough to implement in games on an SSE configured Celeron D (depending on resolution and Bilinear filter & resize

Why ? Because rotating a polygon is harder than subtracting or adding width, Hight & direction to fully complex polygon Fonts & Polygon lines or curves...

The maths is simple enough to implement in games on an SSE configured Celeron D (depending on resolution and Bilinear filter & resize.

Such an example is my SiMD & MMX > AVX Image resizer,
Mipmapping fonts does tend to require over sized fonts..
For example Size 8 & 9 font output = Size 10 to 14 Font,

TT-SVG & Open Fonts OT-SVG & Bitmap fonts compress well;
Mipmapped from 3 sizes larger & Cached as a DOT3/5 or NV12...
You have to save a cache; The Cache can be:

Emulated or Dynamic Spacing (for difficult SETSPACE Console Font situations)
2 Tone, Grey, RGB, RGBA_8888, RGBA_1010102, RGBA_F16, P010, 444A, 888A or 101010A &
(DSC Precached Predicted Block Compression)tm

The representation with alpha is mainly for smoothing & clean lines & is very quick to draw.

Therefore we can Cache a Bitmap Version of any font,
We can of course Vector Render A font & directly to compressed surface rendering.

The full process leads up to the terminal & how to optimize CON,
We can & will need to exceed capacities of any system & To improve them!

*

DSC Precached Predicted Block Compression


We have a font for example with Alpha stored in the screen buffer & of a set size for BLITTING on top of a colour or image background,

The alpha prevents the transposed X-OR Image or Font from having noise & creates ..a smooth sharp in-place modification of content.

For our purpose X-OR can use Alpha instead of a single colour because this allows a very delicate smooth presentation on top of the background..

Repeated application (& Probably Saving of, To save Resource usage); Can overlay graphic of Font Content.

*

VecSR is really good for secondary loading of sprites & text; In these terms very good for pre loading on for example the X86, RISC, AMIGA & Famicom type devices,With appropriate loading into Sprite buffers or Emulated Secondaries (Special Animations) or Font Buffers.

Font Drawing & Vector Render

Although Large TT-SVG & OT-SVG fonts load well in 8MB Ram on the Amiga with Integer & Emulated Float (Library); Traditional Bitmap fonts work well in a Set Size & can resize well if cached & Interpolated &or Bilinear Anti-Alias & sharpened a tiny bit!

presenting: Dev-Con-VectorE²
Fast/dev/CON 3DText & Audio Almost any CPU & GPU ''SiMD & Float/int"
Class VESA Console +

With Console in VecSR you can 3DText & Audio,

VecSR Firmware update 2022 For immediate implementation in all
operating systems & ROM's

Potential is fast & useful.

*
I will put this in print, My 3D & 2D Vector SiMD standard is the thing that i believe will save the most bandwidth on HDMI & DisplayPort Cables & Enable Vector 3D such as Laser Printers & Laser Screens, At the end of the day WE NEED VECTORS : RS
*

https://science.n-helix.com/2022/04/vecsr.html

https://is.gd/Dot5CodecGPU

*

Web graphics & Games : RS : Deep Colour

For the VESA & HDMI Display Standards & Web ICC Protocols

Integer 16Bit R, G, B FFFF,FFFF,FFFF & F16b R, G, B, A FFFFF, FFFFF,FFFFF, FFFF because F16b has 24Bit Integer & 8Bit float components.

I have been thinking more about F16b; B Float with lower precision 8 bit remainder,
We can use it for HSL with Black to White levels (Light, Dark)

5Bit per colour & light & dark as component 4 : R, G, B, A,

Now before this i proposed F16 & F24 & F32 & F64, So what about the advantages of F16b?

So most websites & games use Unsigned Integer F16; F16b is 24Bit Integer with 8 Bit sub pixel colours..
So the float component is mostly usable for games & major colour paint options in CSS Web page markup..

But the Integer 24Bit allows a LOT of colour & we can use the float component in Super Resolution for precise colour additions & in Video as part of the Mpeg decompositions.

Rupert S

These are the main XRGB : RGBA Reference for X,X,X,X
https://drive.google.com/file/d/12vbEy_1e7UCB8nvN3hYg6Ama7HIXnjrF/view?usp=sharing
https://drive.google.com/file/d/1AMR0-ftMQIIC2ONnPc_gTLN31zy-YX4d/view?usp=sharing

Main interpolation references:

Interpolation https://drive.google.com/file/d/1dn0mdYIHsbMsBaqVRIfFkZXJ4xcW_MOA/view?usp=sharing

ICC & FRC https://drive.google.com/file/d/1vKZ5Vvuyaty5XiDQvc6LeSq6n1O3xsDl/view?usp=sharing

FRC Calibration >
FRC_FCPrP(tm):RS (Reference)
https://drive.google.com/file/d/1hEU6D2nv03r3O_C-ZKR_kv6NBxcg1ddR/view?usp=sharing

FRC & AA & Super Sampling (Reference)
https://drive.google.com/file/d/1AMR0-ftMQIIC2ONnPc_gTLN31zy-YX4d/view?usp=sharing
Audio 3D Calibration
https://drive.google.com/file/d/1-wz4VFZGP5Z-1lG0bEe1G2MRTXYIecNh/view?usp=sharing

*

Camera & HDMI & DP Compression Modes


Camera Modes
4:2:1 , 4:2:2 for the 4K Camera : HDR
4:4:4 for the faster 4K Camera : HDR
4:2:1 , 4:2:2 for the faster 8K Camera : HDR

TV Modes

HDMI 1.4 | 4:2:1 , 4:2:2 , 8bit, 10Bit for HD to HD+
HDMI 2 | 4:2:2 , 10Bit, 12Bit HDR 4K
HDMI 2.1 | 4:2:2, 10Bit, 12Bit, 16Bit 4K to 6K/8K..

Example : 5120x2880x 60000Khz-GPixClock-DataRate GRefreshRate-38.365Hz-DBLScan 4:2:2 12Bit

If we had DSC compression modes installed in firmware ...

BEST MODE : Can we upgrade this dynamically to HDMI 2.1 Standards with firmware & DSC Installed

Question is can we implement BEST MODE for our Quality range & Also utilize DSC & Alternative Texture Mode Compression & Dynamic MAX Speed

Yes We Can RS : DSC, ETC, ASTC & DTX Compression for display frames

Yes for Studio recording 4:2:2 mode offers 2x the resolution & 4 extra Bit for the same money as 4:4:4 : 4:2:2 10Bit, 12Bit, 14Bit, 16Bit : Higher Dynamic Contrast & Colour

Examples
https://youtu.be/VCdrB1b7wfc

https://youtu.be/NIsoSA8uO04
https://youtu.be/Suc0OV_9TiA

Render Folder https://bit.ly/VESA_BT

*

ASTC, EAC, DXT, PVRTC & DSC with firmware updated & need to be
included in the standards & firmware.

YCoCg-R


https://en.wikipedia.org/wiki/YCoCg

The screen content coding extensions of the HEVC standard and the VVC standard include an adaptive color transform within the residual coding process that corresponds with switching the coding of RGB video into the YCoCg-R domain.

The use of YCoCg color space to encode RGB video in HEVC screen content coding found large coding gains for lossy video, but minimal gains when using YCoCg-R to losslessly encode video

Yes for Studio recording 4:2:2 mode offers 2x the resolution & 4 extra Bit for the same money as 4:4:4 : 4:2:2 10Bit, 12Bit, 14Bit, 16Bit : Higher Dynamic Contrast & Colour

HDMI 1.4 | 4:2:1 , 4:2:2 , 8bit, 10Bit for HD to HD+
HDMI 2 | 4:2:2 , 10Bit, 12Bit HDR 4K
HDMI 2.1 | 4:2:2, 10Bit, 12Bit, 16Bit 4K to 6K/8K..

Example : 5120x2880x 60000Khz-GPixClock-DataRate
GRefreshRate-38.365Hz-DBLScan 4:2:2 12Bit

https://www.cablematters.com/blog/DisplayPort/hdmi-2-1-vs-displayport-2-0

https://www.cablematters.com/blog/DisplayPort/what-is-display-stream-compression

https://en.wikipedia.org/wiki/YCoCg

https://is.gd/Dot5CodecGPU

*

Things Task Shaders can (c)RS


https://www.phoronix.com/scan.php?page=news_item&px=AMD-RDNA3-More-5.19-Tasks-RADV

Task Shaders can be launched to implement Elliptic & Polygon MESH & thus create:

Things Task Shaders can implement though MESH Shading & Polygons:

(Direct Load of a preform MESH)

Multi-Threaded+
Tundra & fauna
Polygon Fonts
Video Rendering Polygon interpretative interpolation..
Polygon MESH Conceptualised Vector Audio.
X-OR DSC Blank space removal
Polygon math & viewer & Viewer Angle based dynamic MESH Subtraction & Addition..
Close loop Tessellation

OpenCL Group micro tasks
Direct Compute/DirectedCL Group micro tasks
Multi-Threading
*

"Task shader is an optional stage that can run before a Mesh shader in a graphics pipeline. It's a compute-like stage whose primary output is the number of launched mesh shader workgroups (1 task shader workgroup can launch up to 2^22 mesh shader workgroups), and also has an optional payload output which is up to 16K bytes."

**************

Future minimal VSR : fm-VSR : RS

Inference on any device with a C99 compiler
https://pypi.org/project/emlearn/

to run without activating C99; Installs under Python 3.10+
https://github.com/emlearn/emlearn-micropython
https://github.com/emlearn/emlearn-micropython/releases
git clone https://github.com/emlearn/emlearn-micropython

With EmLearn you can compile really tight models of tensors & random forest & Gaussian Matrix,
These are very good for:

A1: Anti-Aliasing ( Gaussian, Tensor error diffusion, forested Random spread )
A2: sharpening & Shaping ( Tensor Edge detect with enhance, Gaussian estimation & line fill, Random forest A to B to D: E to B to F X + )
A3: Line & Curve estimation fills & Tessellation ( forested Random spread (Dither fills) & A1 & A2 & Differentiation in 3D Space : 1:2:3{ A B C : E B F }
A4: HDR & WCG, Combinations of dithering in colour space & light/Shadow differentiation in 3D Space : 1:2:3{ A B C : E B F }

https://science.n-helix.com/2019/06/vulkan-stack.html

VSR https://drive.google.com/file/d/1hewfYqLmY0z-Am800LMR-6H-P5J0Sr0N/view?usp=drive_link

VecSR https://drive.google.com/file/d/1WDvpD9a6TttMTmIz_sRYWaQT3RExBuSq/view?usp=drive_link

https://science.n-helix.com/2022/10/ml.html

https://science.n-helix.com/2021/03/brain-bit-precision-int32-fp32-int16.html

https://science.n-helix.com/2022/04/vecsr.html

https://science.n-helix.com/2016/04/3d-desktop-virtualization.html

https://science.n-helix.com/2022/09/audio-presentation-play.html

Innate Compression, Decompression

https://science.n-helix.com/2022/03/ice-ssrtp.html

https://science.n-helix.com/2022/09/ovccans.html

https://science.n-helix.com/2023/02/smart-compression.html

ML tensor + ONNX Learner libraries & files
Model examples in models folder

https://is.gd/DictionarySortJS
https://is.gd/UpscaleWinDL
https://is.gd/HPC_HIP_CUDA