Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🤝
Open to Collab
Dipankar Sarkar
PRO
dipankarsarkar
7
1657
329
Follow
b4ph's profile picture
WANG-CC's profile picture
chaoliangUNSW's profile picture
127 followers
·
450 following
https://www.dipankar.cc
dipankarsarkar
dipankar
dipankarsarkar
AI & ML interests
Building the AI-native stack. Agents as infrastructure, safety as architecture, performance as plumbing. I publish the receipts: papers, datasets, demos.
Recent Activity
reacted
to
comgen42
's
post
with 🔥
1 day ago
Kodiak v0.4 is out: cortex-agent-llc/kodiak-v0.4-1b (plus an accuracy mode that averages three runs). This release came out of public feedback. After v0.3, someone showed that two of its skills were answering from keywords instead of reading the case. So we built tests a keyword shortcut can't pass: the same case twice, with one detail changed so the right answer flips. Refund eligibility is now fixed: it gets both versions right 79% of the time, up from 22%. The push-to-main-with-failing-tests case that started this now gets "ask the user first". Agent step safety improved but isn't fixed, and the model card says plainly not to use it as a safety control. Sarcasm still shows no signal on real tweets, and accuracy on brand-new kinds of task hasn't moved. That's the big problem for the next version. Now a pause. I'm shutting the training box down for two weeks while I travel in Turkey. I'll be walking through ancient ruins: places where people built things that lasted thousands of years without any of our tools. I'm hoping that does what travel usually does for me, shakes loose some ideas. I'll be thinking about how to teach Kodiak to handle tasks it has never seen, how to run my publishing company better, and where agentic automation actually earns its keep. No training while I'm gone. When I'm back, Kodiak gets the ideas. Every number, including the misses, is in the public build log: github.com/grizzlypeaksoftware/kodiak
reacted
to
appvoid
's
post
with 🔥
1 day ago
Byte-level is surprisingly powerful, don't believe me? have a look at this model https://huggingface.co/dotlabs/void.1
reacted
to
Twu31
's
post
with 🔥
1 day ago
Can a Jev-style decision interface work on EEG when the questions are asked in words? In our tests, only for questions the model was trained on. A small head over frozen EEG features (encode a window once, answer several typed questions about it) was asked each question by a question number, a label template or a description. On SSVEP (BETA, 70 people), asking a seen question with a label template instead of its number cost −5.96 pp of balanced accuracy on a plain spectrum (interval −6.84 to −5.08 pp), the 2 pp margin not met. On flicker frequencies the head was never trained on, it reached 28.7%, where training-free CCA reached 80.9% (chance 12.5%). Results and limits: https://bci.report/topics/questions-in-language/ What "Jev-style" means here: https://bci.report/jev-style/ Query the numbers over MCP: https://huggingface.co/spaces/Twu31/bci-report-explorer Related write-up: https://huggingface.co/blog/Twu31/one-eeg-encoder-several-questions
View all activity
Organizations
dipankarsarkar
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
6 datasets
2 days ago
oro-ai/orobench-trajectories
Viewer
•
Updated
3 days ago
•
5.51k
•
110
•
1
AQ-MedAI/MedMemoryBench
Viewer
•
Updated
2 days ago
•
236k
•
57
•
2
Charly-chan/SWE-Game
Preview
•
Updated
1 day ago
•
620
•
1
mirobody/MetAgentBench
Viewer
•
Updated
3 days ago
•
120
•
290
•
2
Snapkitty/toolgate-bench
Viewer
•
Updated
3 days ago
•
96
•
40
•
1
zlab-princeton/SWEeper-Bench
Viewer
•
Updated
3 days ago
•
200
•
85
•
2
liked
a model
3 days ago
Bayway/JEV-27B-VL-MLX-4bit
Image-Text-to-Text
•
27B
•
Updated
1 day ago
•
418
•
5
liked
a dataset
3 days ago
pbhappliedsystems/quant_eval_paired_degradation_statistics
Viewer
•
Updated
Aug 20
•
48
•
58
•
1
liked
a model
3 days ago
pollix/stuntd-support-triage
Text Classification
•
Updated
3 days ago
•
4
liked
a Space
3 days ago
Running
on
Zero
Agents
2
quant-eval Agent Arena
⚗
2
Compare two LLMs side‑by‑side on any prompt
liked
6 datasets
4 days ago
respanai/guardrail-benchmark
Viewer
•
Updated
12 days ago
•
2k
•
307
•
1
Vineethsain/runtime-guardrail-paper-artifacts
Viewer
•
Updated
8 days ago
•
1
•
296
•
1
harvardMadsys/freeinference_agentic_trace
Viewer
•
Updated
4 days ago
•
2.51M
•
941
•
28
raxITLabs/decision-models-as-guardrails
Viewer
•
Updated
3 days ago
•
9.31k
•
402
•
1
shiweid1/Procedure_Memory_System_Benchmark
Updated
5 days ago
•
155
•
1
Tropic-AI/lumen-bench
Viewer
•
Updated
about 3 hours ago
•
20.6k
•
49
•
1
liked
a model
4 days ago
cortex-agent-llc/kodiak-v0.3-1b
1B
•
Updated
4 days ago
•
3
liked
a dataset
4 days ago
FineEnvs/HF_ML_Tasksmith
RL Environment
•
Updated
16 days ago
•
41.8k
•
4
liked
a dataset
5 days ago
sriram1983007/sra-stablecoin-risk-bench
Viewer
•
Updated
4 days ago
•
40k
•
253
•
2
liked
a model
5 days ago
CountingSheep/vev-4b
Image-Text-to-Text
•
5B
•
Updated
5 days ago
•
172
•
5
Load more